Blog
Notes from the
infrastructure.
Engineering posts, research notes, and customer stories. Published when there's something worth saying — never on a content calendar.
The Pixel Tax
GPT-6 Astra made the screen a universal interface. The bottleneck is no longer connectivity — it is deciding which context deserves compute.
Cosavu Research
Sep 8, 2026 · 14 min
Recent posts
STAN-1-Mini: Training a 12M-parameter RL policy to compress prompts
How we trained a tiny reinforcement learning model to outperform every prompt-compression heuristic on the market — and why 12M parameters was exactly the right size.
The Engram filter: 7 signals, one score, no false positives
Inside CAR-1's hybrid re-ranking layer. Why semantic similarity isn't enough, and what we add on top of it to filter noise from production retrieval.
Why we don't use ChromaDB or FAISS
We started with off-the-shelf vector stores. Three rewrites later, we built our own. Here's what broke and what we learned.
Introducing VexaAgent — production agents you can actually trust
Today we're opening the private beta of VexaAgent, an agent runtime built on ContextAPI, DataAPI, and Vexa-1.
PromptIR: a typed intermediate representation for prompts
Treating prompts like a compiler input. How typed blocks let us optimise without losing intent — and why every block has its own compression strategy.
Sub-5ms inference on CPU — STAN's deployment story
How we got an RL policy network running fast enough that no one notices it's there. ONNX runtime, quantization, and the bottleneck nobody talks about.
How Northwind cut their LLM bill by 47% in three weeks
A short customer story about replacing four separate vendors with Cosavu — and what they did with the saved budget.
Newsletter
Get new posts
in your inbox.
One email per month. Engineering deep-dives, new product announcements, the occasional research note. Never marketing.