Object store is quickly becoming the new core data substrate. Lets build kafka, but on s3. Lets build github, but on s3. It feels like we going to see more and more "object-store first" systems in the next few years.
I am excited about this future. Give me stateless servers and a storage bucket over having to manage systems with disks any day.
I do wonder if we will see an expansion of the s3 api to support more of these use cases. S3 added a janky file append operation to their new express-one-zone bucket type, and limited to 10k total file append operations. I wonder what else we will get in the next few years.
It makes perfect sense: it's a extremely reliable, infinitely-scalable, strongly consistent (in some implementation), cheap, and extremely simple to use key-value store. As such, it transparently solves a lot of the problems distributed systems have to be engineered around.
As long as you can run on a CSP, and can engineer around the high-ish latency (most business cases can), it's extremely expensive to try engineer around it.
The range of things you can do with blob storage and a (very simple) auth model are surprisingly broad.
We recently replaced our Docker container registry with S3 using a tiny tool [1] we built in-house. I think that even with current capabilities, we can still model a lot services as a very thin layer over object storage.
From the opposite side of things, I feel the same about OPFS. Finally having something that performant that multiple workers can operate against in a browser is such a massive boon. You could plug something like that in to it locally with markedly less bullshit than you would have had to do previously with the other APIs.
I'm a little surprised the Docker CLI doesn't support this directly, since gcr.io works this way. Maybe there's a thin layer between the bucket and the CLI? Neat that you worked it out for the generic case.
That sounds really cool — I tend to agree with you that S3 and similar are underutilized, but I remember that essentially all of the providers charge for bandwidth measured in gigabytes and I'm like no. My consumer line is measured in megabits per second and if I have to pay for my data usage the way it's paid for in data centers it would be far more expensive. Somehow consumer ISPs, who have to pay for the lines, are cheaper than than the cloud providers.
"I remember that essentially all of the providers charge for bandwidth measured in gigabytes" - sure, but the cost per GB is shockingly cheap. If you're doing things well, you can do a lot inside those pricing structures.
Alternatively you can stand up your own object store services, but that's not something I would like to do.
> My consumer line is measured in megabits per second and if I have to pay for my data usage the way it's paid for in data centers it would be far more expensive
Well because your provider assumes you are not using all your bandwidth constantly. Cloud bandwidth is only billed for you actually use
Definitely agree that every data system that doesn't need <100ms latency is moving to object storage.
> I do wonder if we will see an expansion of the s3 api to support more of these use cases
This is actually an area where I think we have a big leg up on folks building on top of S3. My team (which built K2) sits next to the R2 team, and we have the opportunity to co-evolve the products in mutually beneficial ways.
You can download objects from S3 from EC2 without traversing the public internet using things like gateway endpoints [0] which avoids s3 egress fees. But doesn't avoid egress fees from EC2 to the end user.
People like S3 because they have hard engineering guarantees around bit rot and work well as a high level abstraction of a network filesystem with all of the low level failure recovery built-in. You don't have to worry about doing your own RAID configs. The bigger question is whether non-AWS services can offer the same level of guarantees. I have heard horror stories for example when it comes to downtime on Hetzner's S3 object store.
I am currently using Cloudflare R2 right now and if you see their forums, there's always the occasional post about objects going missing.
Yes! We originally built K2 to serve as the ingestion layer for Basin Pipelines [0], our stream processing product. We have a number of other teams building new products on top of it at the moment which I can't talk about yet :)
Congrats on the launch! Stream/event-based system are really powerful, but they are also just pretty complex, in no small part because the modeling of streams for most people today is really modeling Kafka topic/partitions which has a whole bunch of foot-guns and complexity. Making the individual stream really cheap and easy is a big simplification, especially if flexibly consuming a stream for both ordered and unordered use-cases is made simple. This looks to work for unordered, be curious to see how the managed dividing the work for the mentioned key-based ordering.
How does this solution differ from AutoMQ and WarpStream? I’ve worked with one of them, and it is indeed a serverless solution built on top of S3.
As far as I know, Kafka itself already supports offloading some data to S3 for long-term storage.
Based on the articles—which I didn't fully grasp—I’m wondering if there are additional benefits mentioned, such as multi-region distribution (though I find it hard to imagine how that would be implemented).
If I was a serious Cloudflare customer I would be seriously concerned about the security of my infrastructure with them. Yes LLMs can code fast but this is an almost frenetic pace of releasing new products, all with fewer staff.
> eventually becoming a monopoly (or part of the big tech "duopoly")
At this point this has already effectively happened. The average person doesn't realize the extent of it, because the products Cloudflare builds are inherently transparent to the average consumer.
I trust that this is not the case. Check out both the authors. Micah's company was acquired by Cloudflare, and both of these guys actually bring good expertise around this area.
The question is less about "vibe coding" the product, and more about how their acquired teams function. For most products they have released so far, they are backed by a company they acqui-hired. They will usually rebrand the product and absorb the team.
You would not be surprised if this release was from a product for a startup company rather than a large enterprise company.
Cloudflare hired a bunch of folks for sure, but they also fired a lot of folks. What they are doing these days is building products by buying out entire companies, giving them nearly independent authority to build a product like they would build a company.
They are working on a set of primitives that make building this kind of software easier. With AI, runtimes matter more than ever and languages matter less.
Yep. I'm not personally a huge fan of the kafka API — I think it's simultaneously too low level for normal users and too high level to deeply integrate into other systems (like stream processing engines), and requires a complex client library to use effectively.
We went with a simpler and more user friendly consume API, that also allows much higher levels of read parallelism (particularly important if you're using something like Workers, which parallelize well but aren't very powerful individually).
But we know many companies are invested in the Kafka ecosystem, and we want to provide an easy on (and if necessary, off) ramp for them.
Google PubSub is a great product. The primary benefit of K2 is cost, particularly for longer retention periods. Being backed by object storage means that we can store data extremely cheaply compared to disk backed solution, and we pass that on in our pricing.
Compared to self-hosted or cloud-hosted Kafka (e.g., Amazon MSK or Confluent), K2 is much cheaper, and fully serverless. There are no clusters to manage or scale, and consistent performance even as you vastly increase the amount of data.
The main downside is produce (and end to end) latency is higher (around 1s p99) than systems that rely on local disk replication, like Kafka.
So it's great if you're trying to move a huge amount of data around, or for use cases where cost is more important than latency.
Having built a hobo version of something similar (serving a minimal subset of the Kafka API on top of CosmoDB): there’s also the hybrid scenario where the cheap serverless streams are used for scalability and broadcast while a low-latency core is maintained on sharply reduced compute resources. An 80/20 approach that saves a lot and, in our case, reduced cluster (mis)management risks at the same time.
In our case BLOB and large document transfers were handled in parallel, merging them together through object storage is a highly appealing package. Great work!
Ha setup and management is costly. K2 relieves you off that by charging you 0.04/0.04/GB read/write and 0.02/GB/month storage. Pretty useful if your volume is not into multi-GB's a month. GKE pub/sub seems pretty costly by comparison.
i was trying to understand why i would use this over their current offerings of queues, and im really sick so my brain isn't working. so i ran it through ai
You need... | Use
-------------------------------------------------------|-------
“Make sure this job gets done” | Queue
retries / dead-letter handling | Queue
delayed jobs | Queue
distribute jobs among workers | Queue
“Record that this event happened” | K2
multiple independent systems reading the same events | K2
replay old events | K2
ordered event streams | K2
Kafka-like architecture | K2
I am excited about this future. Give me stateless servers and a storage bucket over having to manage systems with disks any day.
I do wonder if we will see an expansion of the s3 api to support more of these use cases. S3 added a janky file append operation to their new express-one-zone bucket type, and limited to 10k total file append operations. I wonder what else we will get in the next few years.
As long as you can run on a CSP, and can engineer around the high-ish latency (most business cases can), it's extremely expensive to try engineer around it.
The range of things you can do with blob storage and a (very simple) auth model are surprisingly broad.
We recently replaced our Docker container registry with S3 using a tiny tool [1] we built in-house. I think that even with current capabilities, we can still model a lot services as a very thin layer over object storage.
[1]: https://github.com/Simple-Observability/grue
Alternatively you can stand up your own object store services, but that's not something I would like to do.
Well because your provider assumes you are not using all your bandwidth constantly. Cloud bandwidth is only billed for you actually use
> I do wonder if we will see an expansion of the s3 api to support more of these use cases
This is actually an area where I think we have a big leg up on folks building on top of S3. My team (which built K2) sits next to the R2 team, and we have the opportunity to co-evolve the products in mutually beneficial ways.
[0] https://docs.aws.amazon.com/vpc/latest/privatelink/vpc-endpo...
I am currently using Cloudflare R2 right now and if you see their forums, there's always the occasional post about objects going missing.
Consider your cloud costs.
As someone who enjoys writing Kafka streams applications I am also looking forward to the day you support the Kafka APIs.
Having a cost efficient fully serverless Kafka compatible service would be great, and something I think many businesses would find useful.
Great work!
[0] https://developers.cloudflare.com/basin-pipelines/
As far as I know, Kafka itself already supports offloading some data to S3 for long-term storage.
Based on the articles—which I didn't fully grasp—I’m wondering if there are additional benefits mentioned, such as multi-region distribution (though I find it hard to imagine how that would be implemented).
My bigger concern would be their increasing grip on a lot of the market, and eventually becoming a monopoly (or part of the big tech "duopoly")
At this point this has already effectively happened. The average person doesn't realize the extent of it, because the products Cloudflare builds are inherently transparent to the average consumer.
The question is less about "vibe coding" the product, and more about how their acquired teams function. For most products they have released so far, they are backed by a company they acqui-hired. They will usually rebrand the product and absorb the team.
You would not be surprised if this release was from a product for a startup company rather than a large enterprise company.
Cloudflare hired a bunch of folks for sure, but they also fired a lot of folks. What they are doing these days is building products by buying out entire companies, giving them nearly independent authority to build a product like they would build a company.
There was just more news about it than with other companies ( I think they let go about 2 k. People a year ago).
( Not saying it's good, just a little perception balance)
We went with a simpler and more user friendly consume API, that also allows much higher levels of read parallelism (particularly important if you're using something like Workers, which parallelize well but aren't very powerful individually).
But we know many companies are invested in the Kafka ecosystem, and we want to provide an easy on (and if necessary, off) ramp for them.
Compared to self-hosted or cloud-hosted Kafka (e.g., Amazon MSK or Confluent), K2 is much cheaper, and fully serverless. There are no clusters to manage or scale, and consistent performance even as you vastly increase the amount of data.
The main downside is produce (and end to end) latency is higher (around 1s p99) than systems that rely on local disk replication, like Kafka.
So it's great if you're trying to move a huge amount of data around, or for use cases where cost is more important than latency.
In our case BLOB and large document transfers were handled in parallel, merging them together through object storage is a highly appealing package. Great work!