Tech

Token probability experiment brings language models to webcam image classification

A personal test used vision-capable models to classify webcam frames, with processing rates varying across local and hosted setups.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · View original source
Tech
No image available
Artificial intelligence

A blog post published on 26 September describes adapting a Jev-like language-model wrapper to classify webcam images using token probabilities. The author added an image attachments field to local experiments, sending webcam frames as base64 JPEGs.

The example asks whether a person is visible, whether the scene is indoors or outdoors, and how bright it is. It prompts for a one-token answer and uses the model’s log probabilities for alternative tokens to estimate each classification.

The author reports that Gemma 4 12B running through llama.cpp on an RTX 3090 processed about one frame per second across three questions per frame. OpenAI’s gpt-6-luna processed about 0.2 frames per second in the author’s test.

The post attributes the hosted result partly to separate connections for each question and frame, which the author had not tried to avoid. These are personal test results, not a controlled comparison.

The author says the approach’s appeal is the ability to change classification conditions using plain language, while noting that specialised computer-vision models are likely more efficient. The example also adapts request handling for llama.cpp Chat Completions and OpenAI Responses.

Continue reading

More from Tech

Read next: Amflow’s TL Carbon crosses trail and city, with trade-offs
Read next: Sahai urges investment in people who can scrutinise AI discoveries
Read next: Trump’s FDA nominee backs MMR vaccine but dodges questions on president