From the source
An efficient endpoint for asking one visual asset many questions Sep 4, 2026 Introducing Perceptron Multilook API An efficient endpoint for asking one visual asset many questions Sep 4, 2026 Introducing Perceptron Multilook API An efficient endpoint for asking one visual asset many questions Today we're releasing an endpoint that prefills a shared media context once and reuses it across up to 16 prompts in a single call.
You send video and image context once, attach a list of questions, and get back one result per question.
Multilook costs less, returns sooner, and moves more work through the same fleet.
The cost savings holds on any fleet at any load.
The throughput and latency savings appear when the fleet is busy, which is when they matter.
Reducing the cost of re-sending context The default way to ask several questions about one video is several requests.
The Perceptron Files API took the byte transfer out of that loop: upload once, reference file-abc123 from then on.
The prefill challenge remained unsolved: The model re-reads the media for every question.
A file id saves the upload.
It does not save the forward pass over the frames.
…





