Hi HN, we built an open source model gateway. It's a single place to manage our own self hosted, frontier, and open source models in one place.
It’s is rust native, built for concurrency, and implements all the config quirks across models and providers (streaming formats, tool calls, model parameters, rate limits, and different error behavior).
The gateway adds under 1 ms for BYOK requests and under 2 ms when Experiential supplies the provider key. It has every major inference provider, and 1000+ models refreshed daily via a codex agent that opens a PR.
Compared to other similar projects we’re open source, take no markup, allow you to mix local models with a marketplace, and use your traffic to (opt in) train you a model. Simple routing doesn’t warrant a 10% token markup.
The way we do this is given standardized OTel traces, we mine representative real tasks, use text world models to simulate rollouts for various models, apply an LLM judge, and fit a nearest neighbor classifier on top of an embedding of a prompt to decide the optimal model for each request. Usually this can map out a better pareto curve on cost/quality than just calling single models but it’s not perfect.
Using these simulations we can also do things like suggesting cache hit optimizations, new model suggestions, and training models.
It’s open source, so you can deploy it on your own infrastructure, use our hosted version with 0 markup, or read how we design for maximum availability on our website.
Areibman4 hours ago
Could you say more about how caching works? One major advantage of sticking with a single model is saving money on cached input tokens. I'd imagine if you swap between a bunch of models, you may improve performance but cost would would balloon out of control
SilenNop4 hours ago
The trick is to rarely switch, or switch at task boundaries. Often the conclusion of routing is actually "this one model is actually at the pareto front for this task, just use it always".
cameronh902 hours ago
But then it's better to just not have a gateway switch models at all.
Just have the harness able to choose which model its sub-agents use, then tell it how to split up tasks and which models to use when doing so.
SilenNop2 hours ago
That is another way to do. Or we can automatically figure out which models the subagents should be using for you. And update them as new models come out and the work your subagents do changes. More than one way to skin a cat.
purplecats4 hours ago
and caching is related to performance too ofc
ceroxylonan hour ago
>The gateway adds under 1 ms for BYOK requests
Amazing! Really brilliant idea, thank you for sharing this project. There is so much ground to cover in the LLM gateway / routing / reporting world, and this is a great start. The Tinker implementation is my favorite part, fine tuning is much better than a sea of context files.
kfallah15an hour ago
Thanks! We are going to add continual RL via Tinker soon too
akshay_akulaan hour ago
Open source and no markup is the right default for a gateway. The caching question above is the one I would want answered before swapping models though.
SilenNopan hour ago
Ans: we rarely switch, often times it's just a "switch to using this model for your agent"
cheema332 hours ago
I have not tried it yet. Is it similar to LiteLLM? If so, what sets it apart?
kfallah152 hours ago
Router and model optimization from traffic is the main differentiator
SilenNop2 hours ago
Also a hosted marketplace, not just BYOK
0xbadcafebee38 minutes ago
You started it a week ago? I look forward to checking back in 3 weeks when you've exited for $1B
SilenNop13 minutes ago
See you soon
23david4 hours ago
Super interesting and congrats on the release. Curious if you initially had this in Python and then rewrote in Rust?
SilenNop3 hours ago
Yep! If you look at the commit history that's exactly what happened.
ashermania4 hours ago
Finally an open source tool doing this!
jing09928an hour ago
[dead]