Rendered at 04:50:06 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
creativeSlumber 2 hours ago [-]
> then there are 10^100 possible paths through that session.
Is this correct? in practice you would only be deciding what the next turn is going to be. Because the turn after the next turn is determined by the outcome of the next turn. So shouldn't this number be 10 * 100?
rlprlprlp 2 hours ago [-]
paths friend. so first turn you have ten options, second u have the same ten. assume turn one model h1. turn two you have ten options. 1x10. but you could have chosen any h on the first turn. so 10x10=100. turn three you could have gotten there 100 ways and still have ten options.
adchurch 2 hours ago [-]
My calculation exactly!
YuechenLi 8 hours ago [-]
Have you compared this to using GPT-6.1 Sol instead of GPT 6 Astra + Deepseek? From my test, 6.1 Sol is a lot more token efficient than 6 Sol while being similar to Astra in performance, and I don't really find 6 Astra to be significantly better than 6/6.1 Sol for general coding as I feel 6 Astra is only noticeably better at spatial reasoning/vision compared to 6 Sol, and 6.1 Sol really closed the gap on that front.
conception 5 hours ago [-]
6.1 Sol though is horribly slow. The time cost alone offloading to ds flash is probably worth a look.
YuechenLi 4 hours ago [-]
I don't really mind the speed, I like watching Codex work most of the time, slower work means I have time to do corrective nudging for when the initial prompt was unclear.
thefourthchime 10 hours ago [-]
Interesting work, and thanks for describing how your router works internally. It's definitely a fascinating subject. How would you say this compares to Cursor's auto mode?
adchurch 9 hours ago [-]
Absolutely!
Conceptually very similar to Cursor's auto mode. The key distinctions are:
- We plug into any harness (e.g. Claude Code, Codex, OpenCode, Pi)
- We aren't incentivized to route to our own model, we're incentivized to route to the best model whatever it may be
aschla 10 hours ago [-]
And similarly, Copilot’s Auto mode?
adchurch 7 hours ago [-]
I haven't tried it as recently but last time I checked they only route once per session (or subagent). Imo this is basically impossible to do correctly. Consider the case where you start with one prompt "rewrite this in rust". Trivial in a 1 day old repo, extremely difficult in e.g. the VSCode repo!
361994752 3 hours ago [-]
Sounds very interesting. Do you have plan to support gh copilot as model provider?
We support all the LLMs that Copilot sends data to!
361994752 2 hours ago [-]
Ofc :)
Asking this because we have copilot at work (assuming it is quite common for enterprise)
svnt 7 hours ago [-]
If you are training on data labeled by frontier models, how do you expect to exceed the performance of frontier models, other than in the cost dimension by recognizing simpler problems and routing to cheaper models?
adchurch 7 hours ago [-]
Different frontier models are good at different things! We'll be the ones combining them optimally.
svnt 3 hours ago [-]
I understand the premise, but I believe success depends on you being able to effectively sort problems for other frontier models better than the frontier.
ajspig1 8 hours ago [-]
How do you handle provider variance on OpenRouter for the opensource models? Or do you use your own hosted version to mitigate this?
And for both opensource and closed source, does the router account for provider quality, or catch it when a provider degrades?
adchurch 7 hours ago [-]
For our hosted version we carefully select which providers we use (we don't use OpenRouter). For the self-hosted version though, OpenRouter does make it much easier to get started (at the cost of that variance potentially affecting quality).
Yes for both! We have some logic to put providers on cooldowns and/or deprioritize them.
gitowiec 7 hours ago [-]
Can it route to locally or LAN hosted Qwen or some other open weights model?
adchurch 7 hours ago [-]
Not yet! Something we're interested in experimenting with though, hit me up at andrew@weaveos.com if you have any thoughts here
rirze 9 hours ago [-]
How does this choose which models to use with any arbitrary set of model providers to work from? And why is an openrouter necessary for self-hosting?
adchurch 7 hours ago [-]
To the point of using buckets of models: as long as there's >0 models available in a bucket, and we can order models in order of fit, we're resilient to different sets of models being available. With that said obviously cutting out some models has a much larger effect than others.
OpenRouter isn't strictly necessary but it does make it easier to not have to set up accounts/API keys with several different providers to get started.
jamesforestwest 9 hours ago [-]
How do you define the model buckets, and what happens when a session genuinely needs a model that isn't in the bucket the HMM picked?
adchurch 7 hours ago [-]
You can think of buckets as models with similar capabilities. So for example Deepseek 4.1 Flash will not be in the same bucket as Astra.
The latter is an interesting question! In practice because the session continues, we can see it's going down the wrong path and escalate. Basically no decision we make when routing is entirely unsalvageable (but we do have a performance penalty for every incorrect decision we make so of course we try to avoid it).
dackdel 1 hours ago [-]
what about jev
2 hours ago [-]
devmor 49 minutes ago [-]
I would like to see this (and any model router, frankly) benchmarked against two things, personally:
1) Claude Code's "advisor mode" (nominally, Sonnet 5.5 with a Fable advisor)
2) Copilot's "HydraFusion" model router/advisor combo.
Specifically I would like to see them compared on architecture planning (both human assisted and hands-off with a draft document) and code review, as these are what I have found the most significant improvement on with multi-model systems.
> our initial hypothesis that an ensemble of models can do better than any single model ever could.
To your hypothesis, anecdotally I find both of these offerings to be far superior to any single model for most tasks of any real complexity, and both to have general frustration/failure cases that single models do not. I would not be surprised that any ensemble approach that utilizes more than one single model meets this hypothesis.
Quite frequently I delegate review and restructuring loops to subagents acting as judges/advisors to tell the primary agent if it met the goal it was instructed to. For some workloads, I will even vet every tool call and user-facing output this way.
1minusp 9 hours ago [-]
Does this allow for a predefined budget?
adchurch 7 hours ago [-]
Yes!
redrove 11 hours ago [-]
Is the model you trained available as open weights?
Is this correct? in practice you would only be deciding what the next turn is going to be. Because the turn after the next turn is determined by the outcome of the next turn. So shouldn't this number be 10 * 100?
Conceptually very similar to Cursor's auto mode. The key distinctions are:
- We plug into any harness (e.g. Claude Code, Codex, OpenCode, Pi)
- We aren't incentivized to route to our own model, we're incentivized to route to the best model whatever it may be
And for both opensource and closed source, does the router account for provider quality, or catch it when a provider degrades?
Yes for both! We have some logic to put providers on cooldowns and/or deprioritize them.
OpenRouter isn't strictly necessary but it does make it easier to not have to set up accounts/API keys with several different providers to get started.
The latter is an interesting question! In practice because the session continues, we can see it's going down the wrong path and escalate. Basically no decision we make when routing is entirely unsalvageable (but we do have a performance penalty for every incorrect decision we make so of course we try to avoid it).
1) Claude Code's "advisor mode" (nominally, Sonnet 5.5 with a Fable advisor)
2) Copilot's "HydraFusion" model router/advisor combo.
Specifically I would like to see them compared on architecture planning (both human assisted and hands-off with a draft document) and code review, as these are what I have found the most significant improvement on with multi-model systems.
> our initial hypothesis that an ensemble of models can do better than any single model ever could.
To your hypothesis, anecdotally I find both of these offerings to be far superior to any single model for most tasks of any real complexity, and both to have general frustration/failure cases that single models do not. I would not be surprised that any ensemble approach that utilizes more than one single model meets this hypothesis.
Quite frequently I delegate review and restructuring loops to subagents acting as judges/advisors to tell the primary agent if it met the goal it was instructed to. For some workloads, I will even vet every tool call and user-facing output this way.