Blog
Notes on computer vision, world models, multimodal AI, and research engineering.

Cameras already see everything most sites need. A robot only earns its cost when cameras can't see the information, or can't be installed at all.
TODO

A defense startup event made me realize you can't fake a demo when failure kills people — what if every product were held to that standard?
Once a system can act on what it predicts — not just warn — it stops being a VLM and becomes a world model with a policy on top.
Real proactivity isn't answering better — it's predicting multiple futures, judging which matter, and deciding on its own whether to speak.
An honest look at where continuous latent dynamics genuinely beat discrete-step world models — and where they don't.
Predictive world models and continuous-time dynamics aren't competitors — JEPA learns what matters, Neural ODEs learn how it evolves.