Skip to content
AI.info

The Pulse

Perceptron Releases Mk1.5 for Multimodal Embodied Agents

Perceptron released Mk1.5 on September 25, adding native audio, video tracking, tool use and sub-agent calls to its model for embodied agents. The company says the model works across drones, quadrupeds, smart glasses and phones without plat

Perceptron Releases Mk1.5 for Multimodal Embodied Agents

AI.info Team ·

One model, several kinds of input

Perceptron released Mk1.5 on September 25, describing it as a model built to control embodied agents. The system accepts text, images, video and audio, and can return text alongside structured outputs such as points, boxes, polygons, video clips and object tracks. The release adds native audio, video tracking, web search, sub-agent calls and more complex visual reasoning to the company’s model family.

Perceptron says it has deployed Mk1.5 on drones, robotic dogs, smart glasses and smartphones, and that the model can serve those embodiments without platform-specific retraining. Its approach gives the model a platform’s control surface as a set of tools; Mk1.5 then selects a sequence of calls to complete a task. That design puts the model in an agent loop rather than limiting it to describing an image or video.

Tracking objects across video

For video, Mk1.5 can emit timestamped object tracks instead of returning only separate detections for individual frames. Perceptron says the model led three of the four video object-segmentation benchmarks it measured: Molmo2-Track, Ref-DAVIS17 and ReasonVOS. On MeViS valid_u, the company’s comparison lists MolmoPoint-8B as the leader.

The release also emphasizes first-person video, where camera footage can show what a wearer or robot sees. Perceptron says Mk1.5 localized hands in an egocentric frame 50 percent better than the strongest Gemini model in its comparison. The company links this capability to uses such as annotating robot training footage and helping smart glasses interpret a wearer’s surroundings.

Audio, tools and speed claims

Mk1.5 adds audio input for tasks including time-aligned transcription and sound-based event detection. It can also call functions defined with JSON Schema, and Perceptron says its demo environment includes web search, page reading, zoom and reverse-image search. On the MMSearch benchmark, the company reports a 36.1-point improvement when tool calls are enabled compared with running without them.

Perceptron also reports roughly two to five times faster end-to-end request completion than Mk1, with its latency comparison showing up to 4.7 times faster results. The company says those medians come from three runs on one H100 GPU, with workloads covering chat, image questions and video. The figures are Perceptron’s own measurements; the post does not present them as an independent evaluation.

Available through Perceptron’s platform

The model was available at launch through the Perceptron Platform and an updated SDK. The company lists a 32,000-token multimodal context window and prices of $0.15 per million input tokens and $1.50 per million output tokens. Its API model ID is perceptron-mk1.5.

The release positions Mk1.5 as a perception-and-control component that can combine visual and audio inputs with tools exposed by a device or application. Its benchmark claims describe gains in specific video, egocentric and tool-use tests, while the deployment examples span drones, quadrupeds, glasses and phones. For developers, the concrete offer is a hosted model and SDK that accept multimodal inputs and can issue calls to the control tools they provide.

Source

Explore

More articles