Computer vision
Perception that degrades gracefully — 3D reconstruction, detection and tracking tuned for the conditions a demo never shows you.
Entroπa Labs is an independent AI practice. We take perception, language and robotics from a result that holds on a benchmark to one that holds in the field — and we publish what we learn along the way.
Practice
Benchmarks reward the average case. The world serves the tail — drift, noise, edge conditions. We work where models meet reality, and measure what survives contact.
Perception that degrades gracefully — 3D reconstruction, detection and tracking tuned for the conditions a demo never shows you.
LLM systems you can inspect and trust — routing, evaluation, and the plumbing that turns a capable model into a dependable one.
Control and planning that close the loop under uncertainty — policies that recover when the world refuses to cooperate.
Writing
Not announcements — explanations. Each one takes a single application apart: the mechanism underneath, the part that turned out to be harder than it looked, and the limitations left in rather than tidied away.
7 postsNewest first
every one open to inspection
Our research
Our research ships as software you can run, not screenshots. Each entry below is a live application — the purpose it was built for, the mechanism underneath, and the door into it. Everything executes locally in your browser: no file you open ever leaves your machine.
7 / 7 liveAcross three practices
generative · agentic · physical
The models themselves — opening a checkpoint, reading the weights, deciding which experts earn their parameters.
2 systemsSystems that decide and act unattended — a tool-use loop, a forecast that grades itself nightly, and a decision engine over live data.
3 systemsSystems tied to the physical world — reconstructing a captured space, and driving a body through a simulated one.
2 systemsPurposeJudge a radiance-field reconstruction by moving through it. A 3DGS scene looks flawless from the training views and falls apart two metres to the left; this puts free viewpoint control in front of the reconstruction so the failure shows itself.
PurposeMake a checkpoint legible. A model file is usually a black box you load and hope about. This one opens it in the browser and shows the architecture, the actual weight values, and where the tokens go.
PurposeDecide which experts you can drop, and say what it costs. Most of an MoE's experts are not pulling their weight and several are near-duplicates of each other. This makes that argument with numbers instead of intuition.
PurposeBuild intuition for embodied control with nothing installed. Morphology and terrain decide as much as the controller does; here you can change both in a second and feel the difference on the keyboard.
PurposeRun a probability model in public and grade it every day. Calibration is the only claim that matters and the only one most models never publish. Baseball supplies a fresh, honest test slate every morning.
PurposeTurn public data into the decision a general manager actually faces. Not "who is good" but who is underpriced, who is movable, what a team is missing, and which side wins the trade.
PurposePut the agent loop somewhere you can watch it work. Tool use is usually described in a diagram; here you ask a question out loud and see the model choose a tool, the tool run against real data, and the answer come back — the whole cycle, in the open.
localStorage and sent only to the provider — there is no server here to send it toEach system has a written companion on the blog — what it is for, the mechanism underneath, and the part that was harder than it looked. Papers land here as the work produces them; the shelf is deliberately empty rather than padded. The systems above are the current record: reproducible by construction, offline by default, and honest about which figures are measured and which are estimates.
Contact
Independent AI practice · computer vision, language models and robotics · assessments and hands-on engagements.