NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
How We Pushed CDC into Postgres (snowflake.com)
hasyimibhar 11 minutes ago [-]
It's interesting to watch how different companies that offer both Postgres and warehousing solution under 1 roof approach the same problem:

- ClickHouse focuses on traditional CDC (ClickPipes) and just make it blazingly fast

- Databricks leans on their unified storage architecture (LTAP) to avoid copying data (though you can argue there is still a copy in the cache)

- Snowflake uses a data mirroring CDC as extension so it runs directly on Postgres

I'm still waiting for a Postgres provider to just let me mirror data directly to Iceberg, so I can plug in my own stateless query engine.

bastawhiz 8 hours ago [-]
Clickhouse really nailed this with the acquisition of peerdb. I used it with many terabyte databases and I essentially never thought about it. The only thing we really had to watch for was trying to replicate too much at once (because of the physical compute/io capacity of the postgres or clickhouse clusters).
gopalv 8 hours ago [-]
This was basically Vertica's party trick for quite a long time to have a WOS and ROS formats for the same row and anti-caching between those two.

You could've built a similar system with dezebium and delta lake for quite some time but it would fail compactions, if you run it fast enough. I've seen Oracle GoldenGate 12c do this trick in 2014 or so, using Mysql as the cheap replica. But they are all fragile to schema updates in some direction.

The closest batteries-included equivalent to this is the Aurora -> Redshift bridge[1].

[1] - https://aws.amazon.com/rds/aurora/zero-etl/

bastawhiz 8 hours ago [-]
Aurora zero etl was a nightmare for us. Almost any schema changes require a VACUUM FULL for it to continue functioning. On a few occasions, it just stopped running without an obvious explanation, requiring slow and lengthy back and forth threads with AWS support. If it worked as advertised, it would be great, but I can't recommend it for any serious production system.
fock 8 hours ago [-]
And this is, when it doesn't delete data randomly: https://github.com/trinodb/trino/issues/28885 (some enterprise open source was slop before AI even it seems)
jauntywundrkind 7 hours ago [-]
Although pg_lake is open source, worth noting that it heavily refers to but is missing CDC capabilities.

There's a bunch of comments/links to a closed https://github.com/snowflake-eng/sfpg-extension-pg_lake_repl...

evertheylen 2 hours ago [-]
I'm interested in pg_lake so I wanted to check out your link, but it seems to be internal to snowflake?
plaur782 55 minutes ago [-]
pg_lake is an open source Postgres extension based on work done at Crunchy Data prior to the acquisition by Snowflake - you can find the repo here [1] and a blog post with more context on the project here [2]

[1] https://github.com/Snowflake-Labs/pg_lake

[2] https://www.snowflake.com/en/blog/engineering/pg-lake-postgr...

whateveracct 8 hours ago [-]
Snowflake is a really amazing product. It's been a delight using it the last few years.
p_l 2 hours ago [-]
Seeing Postgres articles from Snowflake surprises me a lot though given how there's zero relation between Snowflake the product and Postgres itself

EDIT: I now see it's mainly to do with pushing data out of customer's postgres systems into snowflake

arvyy 4 hours ago [-]
as much praise as some people give to it, I feel deeply uncomfortable with an idea of SaaS-only DB tech that you don't have an option to self host
8 hours ago [-]
holydementor 5 hours ago [-]
[flagged]
Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 11:20:35 GMT+0000 (Coordinated Universal Time) with Vercel.