Decision Models Launched by OpenAI, Microsoft, Cloudflare, Vercel, Perplexity

· by admin · 6 min read
View as Markdown
Decision Models Launched by OpenAI, Microsoft, Cloudflare, Vercel, Perplexity

Five vendors have launched decision models: OpenAI, Microsoft, Cloudflare, Vercel and Perplexity. A decision model returns typed answers instead of generated text: a probability, a choice from your options, or a score on your rubric. This post sets out what each vendor says about its own model, with links at the end.

The three question types

Each vendor uses slightly different names for the same three question types.

Question type OpenAI Perplexity Vercel
Yes or no predicate noul Boolean
One option from a list choice choice Choice
Rating on ordered levels score score Score

The five launches

Vendor Model Inputs, as stated Input price, as stated Source
OpenAI gpt-6-luna, Decisions API, public beta Text and images $0.10 per 1M input tokens, no output charge OpenAI guide
Microsoft Microsoft-Decision-1 Not stated $0.042 per 1M input tokens, output free Microsoft post
Cloudflare Clef-omni, Clef, Clef-flash Clef-omni: text, images, audio and video Clef-flash $0.038, Clef $0.24, Clef-omni $0.15 per 1M input tokens Cloudflare post
Vercel liquid/d1 from Liquid AI, on AI Gateway Text and images Not stated in the changelog Vercel changelog
Perplexity pplx-decider-v1.1-27b, Decisions API Text, JSON and images $0.02 per 1M input tokens, output free Perplexity docs

What each vendor says about speed

Speed claims come from each vendor’s own comparison, and the baselines differ.

  • OpenAI says the Decisions API is about 10x faster than the Responses API.
  • Microsoft says Microsoft-Decision-1 has a median latency about 35x faster than GPT-6 Sol, and that it is 2.5x faster than H2O-Lightning-4B v1.1.
  • Cloudflare reports a median of about 130 ms for text-only decisions with Clef-omni, about 150 ms for image inputs, and about 1.5 seconds for a 21-second video clip with sound. Its table of serving changes shows the median falling from 262 ms to 152 ms for about 800 input tokens on Clef.
  • Perplexity reports that in its own tests on 30 September 2026, requests with a few hundred input tokens answered in under 2 seconds.
  • Vercel does not state a speed figure in its changelog.

What each vendor says about accuracy

Accuracy claims are the vendors’ own. Each one chose its own benchmarks, so the results can’t be read as a ranking.

  • Microsoft reports the highest accuracy across 36 benchmarks, covering about 150,000 questions, none of which were used in training. It also reports that its decisions changed 1.3% of the time when requests were reworded or options reordered.
  • Cloudflare publishes tables comparing Clef-omni, Clef, Clef-flash and Jev. Its tables show mixed results. Jev scores higher on some benchmarks, such as When2Call, where Jev scores 80.97 and Clef-omni scores 63.3. Clef-omni scores higher on others, such as BANKING77, where it scores 94.8 and Jev scores 79.74.
  • OpenAI, Vercel and Perplexity do not publish accuracy benchmarks in the pages reviewed for this post.

Pricing and context limits

Input prices are listed per million tokens, and each vendor counts tokens differently for images and audio. Cloudflare notes that its conversion rules for image and audio tokens are in its developer docs.

  • Cloudflare says Clef-flash was cut from $0.09 to $0.038 per million input tokens, and that it is now cheaper than Jev. The hosted Clef-flash context window has been cut from 64k to 24k tokens. Cloudflare says only 0.24% of its requests exceed 24k tokens. Clef keeps a 64k context window.
  • Perplexity allows up to 262,144 input tokens per request, with between 1 and 128 questions per request.
  • OpenAI, Microsoft and Vercel do not state a context limit in the pages reviewed.

Input types and request format

The vendors accept different inputs, and their request formats differ.

  • OpenAI takes images only as inline base64 data URLs. It does not accept hosted image URLs or file_id inputs. Its questions are an array, and each question carries a name.
  • Perplexity takes images as base64 data URLs and returns an error for hosted image URLs. Its questions are an object keyed by name.
  • Cloudflare takes text, images, audio (wav or mp3) and video (mp4 or webm) in one call with Clef-omni.
  • Vercel supports image input for liquid/d1, and its changelog shows examples through the AI SDK, an OpenAI-compatible Decisions API, and a TypeSafe-compatible API.

Cloudflare says Clef is fully compatible with the Jev API and can be reached through AI Gateway by changing the model ID.

What this means for an architect

These are recommendations based on the pages above, not claims from the vendors.

  1. Compare on your own workload. The listed prices count tokens differently, and the accuracy results use different benchmarks. Run your own inputs through each candidate.
  2. Match the modality to the decision. If the input includes audio or video, only Clef-omni in this group takes them, according to Cloudflare’s page.
  3. Keep the question set in source control. The request formats differ, so write your questions in one format and convert them in code.
  4. Threshold with margin. Perplexity’s documentation says identical requests usually return identical numbers, but occasionally differ in the second decimal place.

Key Questions

Q1) What is a decision model?

A decision model answers typed questions about an input and returns probabilities, a choice from fixed options, or a score on ordered levels. It does not generate free text. Perplexity’s documentation describes its model as one that “does not write replies, generate code, or explain its reasoning.”

Q2) Which vendor takes audio and video?

According to Cloudflare, Clef-omni takes audio (wav or mp3) and video (mp4 or webm) alongside text and images. The other models named in this post take text and images, as their documentation states.

Q3) Which listed price is lowest?

Based on the input prices each vendor states, Perplexity’s $0.02 per million tokens is the lowest, followed by Cloudflare’s Clef-flash at $0.038 and Microsoft at $0.042. Vercel does not state a price in its changelog. Because the vendors count tokens differently, test cost on your own workload before choosing.

Sources