Perceptron AI just dropped Mk1 (Mark One), a flagship multimodal model aimed at video understanding and embodied reasoning — that's the ability for AI to understand and interact with the physical world. The company has all of 14 employees. It was founded in late 2024 by former Meta FAIR researchers Armen Aghajanyan and Akshat Shrivastava in Washington state, and previously open-sourced the Isaac series of lightweight vision models (2B parameters). Mk1 is their first flagship product.
According to official benchmarks, Mk1 matches or beats Google, Anthropic, OpenAI, and Qwen on image, video, and spatial reasoning tasks — while being drastically cheaper: $0.15 per million input tokens and $1.50 per million output tokens, with a 32K token context window.
The big differentiator is temporal video reasoning. As a hybrid reasoning model, Mk1 can take long videos — like sports games or cooking tutorials — and output a structured timeline analysis, complete with timestamps for specific events. Users can also turn off chain-of-thought when they don't need it to save compute.
On the image side, Mk1 handles pixel-level pointing, dense counting of over a hundred objects, complex OCR, and meter readings. It can also convert complex documents directly into HTML, JSON, or Markdown. These are the kind of tasks that pop up constantly in industrial inspections and warehouse inventory.
For robotics developers, Mk1 treats spatial primitives — points, boxes, polygons, trajectories — as first-class outputs that downstream policy models can consume directly. It can also automatically label teleoperation recordings as training data, cutting out the need for human annotation teams.
Mk1 is available now through the Perceptron API and on OpenRouter.