Cloudflare's Clef: Open Decision Models That Classify Images as Well as Text

On October 1, 2026, Cloudflare released Clef and Clef-flash, two open-source models that answer questions with a typed choice and a probability instead of writing text. Clef accepts images as well as text, so the input to a decision can be a screenshot, a scanned document or a photo, while Cloudflare says Jev only classifies text today. Cloudflare calls these decision models: you give one a state and a list of questions, and it returns an answer and a probability for each. Cloudflare hosts both models on Workers AI and published the weights on Hugging Face under Apache 2.0. It also announced a reinforcement learning service for fine-tuning them on a customer's own data.
Image classification
Cloudflare says Clef has a vision encoder, so it can take in images and classify visual content. Its changelog says you can pass up to four images alongside the state, and Cloudflare's post says Jev only classifies text today.
That makes it possible to classify screenshots, scanned documents, product photos or UI states and get back a typed label with a probability, rather than a caption your code has to parse. Clef isn't the only decision model that takes images. Perplexity's Decisions API also accepts images in the state, per its docs. Strands Decider 2B's Hugging Face card describes a Qwen3.5 text decoder and doesn't mention image input, and Cloudflare says Jev is text-only.
Cloudflare publishes no accuracy numbers for image inputs in the posts we read, so teams should test on their own images before relying on it.
Why it matters
Many of the choices an agent makes are classification problems: which tool to call, whether to retry, whether a message is urgent, whether an output passes a check. A general-purpose language model can answer those, but it generates text first and your code has to parse it. A decision model returns the answer directly, so each call is cheaper, the possible outputs are bounded, and every decision comes with a score you can log and audit.
Jev is a hosted model, and Clef is the open-weight alternative to it. Open weights and an RL fine-tuning service matter to us because a decision model can often be improved by fine-tuning on one team's data, and that's the work we do. We've written about fine-tuning Jev and comparing decision models on a client's own scorecards. The decision models page collects the rest of the measurements.
Key technical notes
- Two models: Clef (27B parameters) and Clef-flash (9B), both built on Qwen backbones, with a 64K-token context window against Jev's 32K, per Cloudflare's post. Clef has a vision encoder and takes up to four images per request, per the changelog.
- The API uses the same System One format as Jev, and the changelog says you can switch an existing Jev integration to Clef by changing the endpoint and model. A request can ask up to 64 questions, in three types: yes/no, multiple choice and scored rubric.
- Cloudflare reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash, against 524.1 ms for Jev. Across 43 benchmark runs, the changelog says Clef is 2.5x faster than Jev at the median and Clef-flash 13x faster.
- Cloudflare reports Clef ahead of Jev on 7 of 10 decision benchmarks, including 94.20% macro-F1 on BANKING77 against 79.74% for Jev. These are Cloudflare's own measurements and we haven't reproduced them.
- The RL fine-tuning service starts with Cloudflare's forward-deployed engineers. Self-serve is planned, using AI Gateway to capture data, Workers AI to roll out models, and a new Trainer component to update weights. The announcement doesn't give pricing.
Sources
- Introducing Clef: our open-source decision models, and new RL fine-tuning platform, Cloudflare blog, October 1, 2026
- Clef on Workers AI, Cloudflare changelog, October 1, 2026
- Cover image: Cloudflare, from the Clef announcement, padded to 1200x630
- TypeSafe AI's Jev now available on AI Gateway, Vercel changelog, September 16, 2026