borkiss*

head of ai · fdaa.dev · ex-quant

--:-- local

* field notes from production

the layer between
frontier models
and production.

i build agent systems that survive production and inference that never gets a gpu bill. what actually works ends up in the two channels below.

everything below is proof. the channels are the point ↓

01 · what's running

a

tiered assistants

production assistants run on three model tiers. cheap models answer instantly, the expensive one wakes up only where money is at stake.

~5% of calls reach the premium tier

b

consilium

sixteen agents argue in isolation, a judge assembles the final answer. the argument itself does the work.

78% of the gain comes from the loop, not the models

c

inference in the browser

llms and video detection run entirely on the user's device. no server, no invoice. 3 live demos ↗

6.5× faster · 724 MB instead of 2.4 GB

50B+tokens a quarter, moving through my agent pipelines
8300+commits a year. my agents write the software
142★bs-p, an avx-512 market-making kernel

02 · the path so far

  1. head of ai, an agency (nda)

    ai strategy, agent pipelines, safety classifiers and the tooling underneath. 50b+ tokens a quarter say it works.

  2. founder & head of engineering, fdaa.dev

    we embed ai engineering into clients' operations. the system stays with them, code and access included. open source on the side.

  3. quantitative trader

    market-making kernels, signals, data pipelines. the adversarial habits from that era are the exact failure modes the agent scene trips over now.

  4. co-founder, holypoly.ca

    cto & cao hat. prediction markets, community, zero boring meetings.

03 · elsewhere

04 · write to me

the fastest way in is telegram.
write here. the rest is noise.