Computer vision
Perception that degrades gracefully — 3D reconstruction, detection and tracking tuned for the conditions a demo never shows you.
Entroπa Labs is an independent AI practice. We take perception, language and robotics from a result that holds on a benchmark to one that holds in the field — and we publish what we learn along the way.
Practice
Benchmarks reward the average case. The world serves the tail — drift, noise, edge conditions. We work where models meet reality, and measure what survives contact.
Perception that degrades gracefully — 3D reconstruction, detection and tracking tuned for the conditions a demo never shows you.
LLM systems you can inspect and trust — routing, evaluation, and the plumbing that turns a capable model into a dependable one.
Control and planning that close the loop under uncertainty — policies that recover when the world refuses to cooperate.
Method
A result is a hypothesis until it holds under load. We instrument everything, chase the failure modes rather than the headline metric, and ship the version that keeps working when the distribution moves. Then we write down why.
Our research
Our research ships as software you can run, not screenshots. Each entry below is a live application — the purpose it was built for, the mechanism underneath, and the door into it. Everything executes locally in your browser: no file you open ever leaves your machine.
7 / 7 livePerception · interpretability · efficiency
embodiment · forecasting · decision support · agents
PurposeJudge a radiance-field reconstruction by moving through it. A 3DGS scene looks flawless from the training views and falls apart two metres to the left; this puts free viewpoint control in front of the reconstruction so the failure shows itself.
PurposeMake a checkpoint legible. A model file is usually a black box you load and hope about. This one opens it in the browser and shows the architecture, the actual weight values, and where the tokens go.
PurposeDecide which experts you can drop, and say what it costs. Most of an MoE's experts are not pulling their weight and several are near-duplicates of each other. This makes that argument with numbers instead of intuition.
PurposeBuild intuition for embodied control with nothing installed. Morphology and terrain decide as much as the controller does; here you can change both in a second and feel the difference on the keyboard.
PurposeRun a probability model in public and grade it every day. Calibration is the only claim that matters and the only one most models never publish. Baseball supplies a fresh, honest test slate every morning.
PurposeTurn public data into the decision a general manager actually faces. Not "who is good" but who is underpriced, who is movable, what a team is missing, and which side wins the trade.
PurposePut the agent loop somewhere you can watch it work. Tool use is usually described in a diagram; here you ask a question out loud and see the model choose a tool, the tool run against real data, and the answer come back — the whole cycle, in the open.
localStorage and sent only to the provider — there is no server here to send it toWritten notes and papers land here as the work produces them — the shelf is deliberately empty rather than padded. The systems above are the current record: reproducible by construction, offline by default, and honest about which figures are measured and which are estimates.
Contact
Independent AI practice · computer vision, language models and robotics · assessments and hands-on engagements.