Private AI, on your own machine — deep enough to know you, fast enough to talk to.
AI is about to fit on the machine you already own. This is the assistant built for that moment.
Why now
Your assistant lives in someone else’s data centre because, until recently, that was the only place a good model would fit. That has an expiry date.
The models and the machines are coming whether anyone prepares or not; what sits on top of them is a choice. Privacy stops being a sacrifice and becomes a side effect. And when the silicon makes running a model trivial, the model stops being the hard part — what it knows about you becomes the whole game.
The models, coming down
Small enough to fit
More capable per gigabyte every release, and published rather than rented. Anyone can download one and run it.
The hardware, coming up
Fast enough to talk to
The machine on your desk keeps getting better at exactly this. When the lines cross, the model never has to leave the room.
It’s running
Ask it something and a small face glances from page to page as it searches — opening one, then the next — then answers out loud in about a second and a half. No spinner. You can see what it read.
And it would rather say I don’t know than invent something. The reason under every answer is the real list of pages it opened, not a story told afterwards.
Under the hood
Deep models understand you but run at about a word a second. Fast models hold a conversation but can’t understand you properly. Kaineros runs both: a slow one that compiles between conversations, and a small quick one that only ever reads finished pages — so the half you talk to never pays for the half that thinks.
While you’re not looking, the slow half argues with itself about what’s true. Rival versions of a fact go head to head and the winner gets written down. An offhand remark has to beat a settled one, not merely be the most recent thing you said.
Its memory is a folder you can read. One Markdown page per thing it believes about you, with where it came from and when. Open it, disagree with it, edit it. Nothing to export, because nothing was ever encoded.
The full argument is in the white paper, and how it compares to everyone else building AI memory.
Get involved
I’m Mat, a UK-based founder, building this in the open. If you work on local inference, open weights or memory — or you just think this is where everything is heading — I’d like to talk. No pitch and no role attached. I want to know who else is out here.
The prototype works: the two-speed engine, a test suite, a benchmark corpus. The repository is private while the project is young — ask, and I’ll open it up to you.
Building in this space as a company? I’m open to working on it — say hello.