tiered assistants
production assistants run on three model tiers. cheap models answer instantly, the expensive one wakes up only where money is at stake.
~5% of calls reach the premium tier
--:-- local
* field notes from production
i build agent systems that survive production and inference that never gets a gpu bill. what actually works ends up in the two channels below.
everything below is proof. the channels are the point ↓
production assistants run on three model tiers. cheap models answer instantly, the expensive one wakes up only where money is at stake.
~5% of calls reach the premium tier
sixteen agents argue in isolation, a judge assembles the final answer. the argument itself does the work.
78% of the gain comes from the loop, not the models
llms and video detection run entirely on the user's device. no server, no invoice. 3 live demos ↗
6.5× faster · 724 MB instead of 2.4 GB
ai strategy, agent pipelines, safety classifiers and the tooling underneath. 50b+ tokens a quarter say it works.
we embed ai engineering into clients' operations. the system stays with them, code and access included. open source on the side.
market-making kernels, signals, data pipelines. the adversarial habits from that era are the exact failure modes the agent scene trips over now.
cto & cao hat. prediction markets, community, zero boring meetings.
the fastest way in is telegram.
write here. the rest is noise.