Hacker News

adesh_nalpet
Show HN: PicoMQ – Durable Streams over HTTP, on object storage picomq.com

PicoMQ is a Rust server for Durable Streams, built on Object Store. Cheap, URL-addressable, granular streams (create/append/read/long-poll/SSE), with Pico Protocol or Durable Streams Protocol as the facade.

S3Stream is the stream storage primitive, used in AutoMQ, shipped as a Rust library. Coordination is a command log in Postgres.


gpsmsn7 hours ago

Love the landing page design! I'm excited to give it a whirl this evening.

Initially, I started to compare PicoMQ to https://github.com/s2-streamstore/s2. If you're familiar with them, how would you compare PicoMQ to S2?

adesh_nalpetop6 hours ago

Thank you! I'd say pretty close. But S2 is not open source, and the OSS version (https://github.com/s2-streamstore/s2) is S2-Lite, which is single-node only and uses SlateDB.

SlateDB is great, but it's not purpose-built for streaming workloads. With PicoMQ, the goal is to really do one thing well. It would have been far easier to just extend SlateDB, but instead, the S3Stream engine is built from the ground up with AutoMQ's core primitives.

Not to mention, Pico supports multi-node/cluster deployments and a bunch of other features (no feature gating).

raiyanr22 minutes ago

PicoMQ looks worth experimenting with during your prototype phase, especially because it is designed to run as a single binary and supports a relatively simple HTTP-based model

adesh_nalpetop7 minutes ago

Indeed! The binary even embeds the admin dashboard. A good starting point to get the hang of PicoMQ: https://picomq.com/docs/playground

Jonovono6 hours ago

I’ve been obsessed with everything durable lately and this is good timing as electric just got acquired so the state of their durable streams project is unknown.

Have you looked into the semantics of their like StreamDB stuff ? Pretty interesting

Also there is this that was built on top of it. Sounds somewhat similar goal as yours but I could be off. Need to explore more https://ursula.tonbo.io/

adesh_nalpetop5 hours ago

On the same boat, and honestly, it’s not getting the attention it deserves. Given Databricks acquired ElectricSQL primarily for PGlite, the Durable Streams project is likely going to be abandoned unless new maintainers step in.

I did, in fact, initially write PicoMQ with OpenRaft, similar to Ursala, but I really wanted the operational complexity to be minimal and the nodes to be stateless (at least for most use cases), Pico uses SQL database as a metadata command log, inspired by RisingWave.

But Ursala would, without a doubt, have better durability ACK latency, as it wouldn’t have to wait for an ACK from S3. That said, I do plan on extending disk/EBS-staged WAL to hit similar low-latency durability ACK numbers. But again, most use cases don’t need single-digit-millisecond latency for ACKs.

alasano3 hours ago

I was planning to use electric cloud for durable streams recently and found out they are winding it down post acquisition.

Unfortunate timing and self hosting isn't complex or anything but still as you said, durable streams as a project is most likely going to be abandoned.

adesh_nalpetop3 hours ago

True, there are also a lot of open PRs and issues. I'm surprised no other developers or orgs are stepping in to keep the project afloat. No wonder the Apache Way goes a long way! Either way, their approach of building around an open-protocol was a good decision.

If you'd like to consider using PicoMQ, I have examples to deploy on Fly.io and AWS. ursula.tonbo.io is also a good option to try.

alasano2 hours ago

[dead]

Jonovono5 hours ago

Thanks for the reply, sounds great, going to dive in!

adesh_nalpetop5 hours ago

Your welcome! And feel free to ask any questions as try out PicoMQ or raise on Github issues. The disk-staged WAL is currently tracked here: https://github.com/PicoMQ/picomq/issues/13, I'm planning to ship it as an optional add-on.

KaiserPro8 hours ago

Can you help an old man understand?

This sounds like a kafka-like streaming system, but backed onto s3-objects?

Doesn't this mean that write performance is going to be bad?

adesh_nalpetop8 hours ago

I'm glad you asked, and yes, PicoMQ does have some Kafka-like semantics. However, Kafka is great at being a huge pipe, so you'd create topics like tables. PicoMQ, on the other hand, recommends creating granular streams that make the most sense, say, by user, session, or vehicle (still bottomless).

And it's also fair to question write performance, since it's backed by object storage. The optimization is primarily from the shared WAL across streams, server-side batching, and client-side in-memory pipelining, especially with HTTP/2, without as much connection pool overhead.

In practice, you can go to the extent of achieving up to 100 MiB/s throughput per stream. Considering how granular streams can be, you'd rarely need as much. The latency for a durability ACK is, however, the price to pay, which is going to be ~250 ms, or lower with S3 Express, which I'd say covers most real-time use-cases. The design itself is easy enough to extend to a disk-staged WAL for single-digit durability ACK latency.

atombender4 hours ago

Have you look at Google's new Rapid Bucket offering? Google Cloud Storage is arguably as good as S3, and is protocol compatible; Rapid Buckets are a type of bucket which supports appendable objects and low latency I/O. The downside is they can only be zonal, and they're a bit more expensive.

adesh_nalpetop4 hours ago

PicoMQ works with any S3-compatible object store. But I wasn't aware of GCS Rapid Bucket, it sounds a lot like AWS S3 Express, which is also zonal. And it does help with durability ACK latency quite a bit, keeping it closer to ~50ms.

I'll be setting up a GCP deployment example similar to AWS soon. I'll be sure to try Rapid Bucket as well, thanks for sharing!

Shakahs8 hours ago

Not familiar with this particular library, but similar libraries use S3 Express One Zone which has write latency <10ms, so you can use that for the WAL and compaction can move data onto other storage classes in the background.

Regular S3 has write latency 100-150ms, which might be fine depending on your workload anyway.

adesh_nalpetop8 hours ago

You nailed it! With some of the similar products I've seen, they either inherit the Kafka protocol and hence the KRaft and other complexities, or go the other way with single-node only deployments, commonly just using SlateDB's single-writer model for durability.

soleveloper8 hours ago

Sounds great, and documentation is very clear.

So - in theory - something like a massive chat client, discord like, can be implemented via this solution? And what would be the pricing of such a solution. Cheap-serverless-discord

adesh_nalpetop8 hours ago

Thank you! I spent a decent chunk of time designing PicoMQ, so the documentation had a natural progression.

Exactly, streams can essentially be rooms, and since the ordering is preserved, a Discord-like application is a strong use case. I’m even considering building one using PicoMQ as an example showcase.

The pricing is going to be dirt cheap, and the best part is how easy it is to scale up vertically or add nodes. For some raw numbers, assuming 1M messages/day, 200B per message, ~6GB/month, all-inclusive, it would be $30 to $150 a month, and storage would be the cheapest part.

soleveloper8 hours ago

Yes, discord as an example would be awesome, exactly what I thought.

Is this back of envelope pricing include the traffic/bandwidth of the readers?

That could easily be 10-100x of number of messages.

This solution together with a cheap/free caching layer (especially for non members/writers) could be amazing.

BTW, a classic example would be an hn mirror. ;)

adesh_nalpetop7 hours ago

Yeah, much of the base cost I mentioned was from the compute, networking and S3 write costs, which could sustain a lot more messages for sure. It also depends on the number of active streams/rooms, the throughput bursts, as opposed to a steady state.

But as we speak, I am in the process of running open-benchmarks on AWS with cost attribution. I'll be posting them in the docs with transparency soon enough.

Read and write through cache already exists today, which would work in favour of both cost, latency, and consumer reads fanout.

I like the HN Mirror over a Discord-like app for simplicity. No doubt that's where my weekend is going!

nulltrace4 hours ago

[dead]

quietfold30 minutes ago

[flagged]

kimseungyong2 hours ago

[dead]

useiris8 hours ago

[flagged]

hathym7 hours ago

[dead]

hn-front (c) 2024 voximity
source