Data · Jev's first week · 16 to 23 September 2026

A week of Jev, sorted

I sorted 5,950 posts from Jev's first week, read the usage the gateways show, and counted the rivals. This is what the data says. What I make of it is a separate piece.

Jev is TypeSafe AI's decision model. It doesn't write text. You give it a state and a typed question, pick one of these, score this, yes or no, and it answers in a few hundred milliseconds with a probability. TypeSafe's launch claim was 20 to 200x faster and 40 to 400x cheaper than an LLM for that kind of question; "100x cheaper, 10x faster" is the shorthand the week settled on.

I built Pong on it the week it came out. The post did well, and nearly every comment asked the same thing: fine, you built a game, what's it for? I tried to answer and couldn't. Every use case I reached for was either already solved or harder than a cheaper yes-or-no. So I stopped guessing and measured: every post in OpenChamber's Jev feed for the first week, sorted and audited; the two gateways that publish usage; and the rivals that turned up. The method and the audit are at the end.

The short version

  • Jev is being called at scale, and the money is small. On 23 September it was a quarter of all requests on Vercel's AI Gateway and 2% of its tokens; on OpenRouter it served more requests that week than GPT-5.6 Luna, for an eighth of the spend. TypeSafe publishes no usage number, and 96% of OpenRouter's Jev requests come from apps that don't say who they are.
  • The buzz is about using it, not about what it makes possible. Four in five posts compared Jev with nothing, and none of the 91 builds most likely to show something new did something that was unavailable before.
  • TypeSafe announced 40 to 400x cheaper than frontier models, and against frontier models it holds. The posts claimed 28x. Measured against the small models you'd actually use, the median is 7.3x on cost and 4.7x on latency, often with better accuracy.
  • Twenty rivals in a week, mostly built by one person in days on open weights. None matches Jev's mix yet, and no big provider has shipped one.

1 · How big it really is

A quarter of Vercel's gateway requests, 2% of its tokens, and nobody can say who's behind it

TypeSafe hasn't published a usage number, so the gateways are the only view anyone has, and they show what the posts can't: Jev is being called at a scale that dwarfs the demos. What the chart can't show is who's calling it. Five named apps (a data-labeling app with 12.9 million requests, Magnific, mirasim, the Asian Development Bank's evidence portal, a Naver rubric judge) account for 4% of OpenRouter's Jev requests. The rest is anonymous. Demand outran TypeSafe: 140,000 people were let off the waitlist in 36 hours, signups were paused a week in, and the API had four short outages.

Jev became a quarter of Vercel's gateway requests in a week, and 2% of its tokensOpenRouter requests a day, 18 to 23 Sep; Vercel AI Gateway share of all requests and of all tokens, 17 to 23 Sep OpenRouter served 126M Jev requests on 23 Sep. Jev was 25% of Vercel AI Gateway requests and 1.8% of its tokens that day.Jev became a quarter of Vercel's gateway requests in aweek, and 2% of its tokensOpenRouter requests a day, 18 to 23 Sep; Vercel AI Gateway share of all requests and of alltokens, 17 to 23 SepOpenRouter: Jev requests a day, in millions0M50M100M150M18 Sep: 28M28M18 Sep19 Sep: 53M53M19 Sep20 Sep: 64M64M20 Sep21 Sep: 82M82M21 Sep22 Sep: 84M84M22 Sep23 Sep: 126M126M23 SepRequests are API calls, not decisions; one call can carry several questions.Vercel AI Gateway: Jev's share of all requests, and of all tokens0%10%20%30%17 Sep: 5.9% of requests18 Sep: 13.6% of requests19 Sep: 26.3% of requests20 Sep: 20.8% of requests21 Sep: 21.2% of requests22 Sep: 22.0% of requests23 Sep: 25.2% of requests25% of requests17 Sep: 0.3% of tokens18 Sep: 0.8% of tokens19 Sep: 1.8% of tokens20 Sep: 2.0% of tokens21 Sep: 2.0% of tokens22 Sep: 1.8% of tokens23 Sep: 1.8% of tokens1.8% of tokens17 Sep18 Sep19 Sep20 Sep21 Sep22 Sep23 SepJev was free on Vercel until 25 Sep, which flatters its share there. Vercel publishes shares, notcounts.Source: OpenRouter and Vercel AI Gateway, read 24 Sep 2026. Method: see the end of this page.
Jev became a quarter of Vercel's gateway requests in a week, and 2% of its tokensOpenRouter requests a day, 18 to 23 Sep; Vercel AI Gateway share of all requests and of all tokens, 17 to 23 Sep OpenRouter served 126M Jev requests on 23 Sep. Jev was 25% of Vercel AI Gateway requests and 1.8% of its tokens that day.Jev became a quarter of Vercel'sgateway requests in a week, and 2%of its tokensOpenRouter requests a day, 18 to 23 Sep; Vercel AIGateway share of all requests and of all tokens, 17to 23 SepOpenRouter: Jev requests a day, in millions0M50M100M150M18 Sep: 28M28M18 Sep19 Sep: 53M53M19 Sep20 Sep: 64M64M20 Sep21 Sep: 82M82M21 Sep22 Sep: 84M84M22 Sep23 Sep: 126M126M23 SepRequests are API calls, not decisions; one call can carryseveral questions.Vercel AI Gateway: Jev's share of requests and oftokens0%10%20%30%17 Sep: 5.9% of requests18 Sep: 13.6% of requests19 Sep: 26.3% of requests20 Sep: 20.8% of requests21 Sep: 21.2% of requests22 Sep: 22.0% of requests23 Sep: 25.2% of requests25% requests17 Sep: 0.3% of tokens18 Sep: 0.8% of tokens19 Sep: 1.8% of tokens20 Sep: 2.0% of tokens21 Sep: 2.0% of tokens22 Sep: 1.8% of tokens23 Sep: 1.8% of tokens1.8% tokens17 Sep19 Sep21 Sep23 SepJev was free on Vercel until 25 Sep, whichflatters its share there. Vercel publishesshares, not counts.Source: OpenRouter and Vercel AI Gateway, read 24 Sep2026. Method: see the end of this page.
Show the numbers
DayOpenRouter requestsPrompt tokensBilledVercel: share of requestsVercel: share of tokens
17 Sep5.9%0.34%
18 Sep28,243,33873B$3,06913.6%0.77%
19 Sep52,641,640140B$5,85226.3%1.79%
20 Sep64,434,110196B$8,16920.8%2.00%
21 Sep82,379,677256B$10,72321.2%2.04%
22 Sep83,638,877246B$10,29522.0%1.75%
23 Sep125,647,722327B$13,65525.2%1.81%

Who is trying it

SignalNumberWhenSource
People let off the waitlist140,000 in under 36 hours17 SepDiogo Almeida, TypeSafe
Waitlist removed"Jev is now available to everyone."20 SepTypeSafe
Signups paused"an immense swell of demand"; not reopened by 24 Sep22 SepTypeSafe
Vercel AI Gateway, day onenearly 13% of paid teams by hour 24; every other recent launch under 7% after a day18 SepVercel
Vercel AI Gateway, latest dayused by 37.8% of teams, first; the primary model by tokens for 33.8% of teamsread 24 SepVercel leaderboards
SDK downloads, launch week298,412 JavaScript, 46,177 Vercel provider, 331,517 Python (automated installs included)read 24 Sepnpm, pypistats
New GitHub repositories matching "jev"8,106 since 15 Sep, against 56 the week beforeread 24 SepGitHub search
Discord members107,727read 24 SepDiscord
Named apps on OpenRouterfive apps hold 4.3% of Jev requests; a data-labeling app alone 12.9 million; the rest is anonymousread 24 SepOpenRouter
IncidentsAPI down 5, 18, 2 and 12 minutes on 17, 20, 21 and 23 Sep; none in July or Augustread 24 SepTypeSafe status page

The calls are real and tiny. A request isn't a decision, a gateway can't tell trying from using, and Jev was free on Vercel all week, so read the Vercel line as a ceiling.

2 · Weighted by money

More requests than GPT-5.6 Luna, for an eighth of the money

This is what 100x cheaper looks like on a bill. Priced on the same day's tokens, Jev sits with the cheapest small models, and one of them, DeepSeek V4 Flash, is cheaper per token. The 244x is only there if you assume the calls would have gone to a frontier model, and nobody sends yes-or-no questions to Fable. Most of these calls exist because they cost almost nothing, not because they replaced something expensive.

More requests than GPT-5.6 Luna, for an eighth of the moneyOpenRouter, week to 23 Sep: Jev against GPT-5.6 Luna; then one day's Jev tokens priced on other models Jev 437M requests, 1.39T tokens, $51,763; GPT-5.6 Luna 315M, 8.74T, $415,706. On 23 Sep Jev's 327B prompt tokens cost $13,655.More requests than GPT-5.6 Luna, for an eighth of themoneyOpenRouter, week to 23 Sep: Jev against GPT-5.6 Luna; then one day's Jev tokens priced onother modelsWeek to 23 Sep on OpenRouter: Jev (blue) against GPT-5.6 Luna (grey)RequestsJev is 1.4xJev: 437M437MGPT-5.6 Luna: 315M315MTokensJev is 16%Jev: 1.39T1.39TGPT-5.6 Luna: 8.74T8.74TSpendJev is 12%Jev: $51.8k$51.8kGPT-5.6 Luna: $416k$416kWhat 23 Sep's Jev tokens (327B prompt tokens) would cost on other models, at list price$10k$100k$1M$10MDeepSeek V4 Flash 0731DeepSeek V4 Flash 0731: $10,209, 0.7x Jev's bill$10.2k · 0.7xJevJev: $13,655, 1.0x Jev's bill$13.7k · billedGPT-5 NanoGPT-5 Nano: $16,848, 1.2x Jev's bill$16.8k · 1.2xMinistral 3BMinistral 3B: $32,816, 2.4x Jev's bill$32.8k · 2.4xGemini 3.5 Flash-LiteGemini 3.5 Flash-Lite: $101,211, 7.4x Jev's bill$101k · 7.4xClaude Haiku 4.5Claude Haiku 4.5: $333,183, 24.4x Jev's bill$333k · 24xGPT-5.6 SolGPT-5.6 Sol: $666,365, 48.8x Jev's bill$666k · 49xClaude Fable 5.1Claude Fable 5.1: $3,331,826, 244.0x Jev's bill$3.3M · 244xSame prompt tokens, plus 10 output tokens a request for the chat models: a floor formodels that think before answering. A price comparison only; nobody has run these taskson those models.Source: OpenRouter rankings and list prices, read 24 Sep 2026. Method: see the end of this page.
More requests than GPT-5.6 Luna, for an eighth of the moneyOpenRouter, week to 23 Sep: Jev against GPT-5.6 Luna; then one day's Jev tokens priced on other models Jev 437M requests, 1.39T tokens, $51,763; GPT-5.6 Luna 315M, 8.74T, $415,706. On 23 Sep Jev's 327B prompt tokens cost $13,655.More requests than GPT-5.6 Luna, foran eighth of the moneyOpenRouter, week to 23 Sep: Jev against GPT-5.6Luna; then one day's Jev tokens priced on othermodelsWeek to 23 Sep on OpenRouter: Jev (blue) againstGPT-5.6 Luna (grey)RequestsJev is 1.4xJev: 437M437MGPT-5.6 Luna: 315M315MTokensJev is 16%Jev: 1.39T1.39TGPT-5.6 Luna: 8.74T8.74TSpendJev is 12%Jev: $51.8k$51.8kGPT-5.6 Luna: $416k$416kWhat 23 Sep's Jev tokens (327B prompt tokens)would cost on other models, at list price$10k$100k$1M$10MDeepSeek V4 FlashDeepSeek V4 Flash 0731: $10,209, 0.7x Jev's bill$10.2k · 0.7xJevJev: $13,655, 1.0x Jev's bill$13.7k · billedGPT-5 NanoGPT-5 Nano: $16,848, 1.2x Jev's bill$16.8k · 1.2xMinistral 3BMinistral 3B: $32,816, 2.4x Jev's bill$32.8k · 2.4xGemini Flash-LiteGemini 3.5 Flash-Lite: $101,211, 7.4x Jev's bill$101k · 7.4xHaiku 4.5Claude Haiku 4.5: $333,183, 24.4x Jev's bill$333k · 24xGPT-5.6 SolGPT-5.6 Sol: $666,365, 48.8x Jev's bill$666k · 49xFable 5.1Claude Fable 5.1: $3,331,826, 244.0x Jev's bill$3.3M · 244xSame prompt tokens, plus 10 output tokensa request for the chat models: a floor formodels that think before answering. A pricecomparison only; nobody has run thesetasks on those models.Source: OpenRouter rankings and list prices, read 24 Sep2026. Method: see the end of this page.
Show the numbers
ModelPrompt price, $ per million tokensOutput priceCost of 23 Sep's Jev tokensMultiple of Jev's bill
DeepSeek V4 Flash 07310.0300.32$10,2090.7x
Jev0.0420.00$13,655billed
GPT-5 Nano0.0500.40$16,8481.2x
Ministral 3B0.1000.10$32,8162.4x
Gemini 3.5 Flash-Lite0.3002.50$101,2117.4x
Claude Haiku 4.51.0005.00$333,18324x
GPT-5.6 Sol2.00010.00$666,36549x
Claude Fable 5.110.00050.00$3,331,826244x

A price comparison, not a quality one: nobody has run these tasks on the other models.

3 · What people built

We can't see what everyone runs. We can see what they show.

The gateways show volume with no names. The posts show names with no volume. So the rest of this page is about the 5,595 posts: what people said Jev decides, what they measured, and what they compared it with. Two things stood out. Games are the biggest single family, not classification. And a fifth of the posts never say what the decision is at all.

A third of the posts are decisions about data, and a fifth are games5,595 posts in six groups by what the decision is about, shaded by what each post measured Decisions about data: 1,936 posts, 35%. Decisions inside games and control loops: 1,071 posts, 19%. Decisions inside agent loops: 603 posts, 11%. Decisions about what a person just said or typed: 463 posts, 8%. Decisions about money: 268 posts, 5%. No decision stated: 1,254 posts, 22%.A third of the posts are decisions about data, and a fifthare games5,595 posts in six groups by what the decision is about, shaded by what each post measuredNo measurementMeasured demoLabeled production, not confirmedProduction, confirmed by the auditDecisions about dataclassify, route, extract, rerank, judge, telemetryDecisions about data, no measurement: 1,135 postsDecisions about data, measured demo: 778 postsDecisions about data, labeled production, not confirmed: 19 postsDecisions about data, production, confirmed by the audit: 4 posts1,936 posts· 35% · 41% measuredDecisions inside games and control loopsgames, simulations, robotsDecisions inside games and control loops, no measurement: 791 postsDecisions inside games and control loops, measured demo: 280 posts1,071 posts· 19% · 26% measuredDecisions inside agent loopsgate a tool call, pick the next browser action, prune contextDecisions inside agent loops, no measurement: 394 postsDecisions inside agent loops, measured demo: 207 postsDecisions inside agent loops, labeled production, not confirmed: 1 postsDecisions inside agent loops, production, confirmed by the audit: 1 posts603 posts· 11% · 35% measuredDecisions about what a person just said or typedmoderation, voice, live chat, collaborationDecisions about what a person just said or typed, no measurement: 338 postsDecisions about what a person just said or typed, measured demo: 121 postsDecisions about what a person just said or typed, labeled production, not confirmed: 4 posts463 posts· 8% · 27% measuredDecisions about moneytrading botsDecisions about money, no measurement: 145 postsDecisions about money, measured demo: 122 postsDecisions about money, labeled production, not confirmed: 1 posts268 posts· 5% · 46% measuredNo decision statedmeta, benchmarks of the model, wrappersNo decision stated, no measurement: 1,032 postsNo decision stated, measured demo: 215 postsNo decision stated, labeled production, not confirmed: 4 postsNo decision stated, production, confirmed by the audit: 3 posts1,254 posts· 22% · 18% measuredSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
A third of the posts are decisions about data, and a fifth are games5,595 posts in six groups by what the decision is about, shaded by what each post measured Decisions about data: 1,936 posts, 35%. Decisions inside games and control loops: 1,071 posts, 19%. Decisions inside agent loops: 603 posts, 11%. Decisions about what a person just said or typed: 463 posts, 8%. Decisions about money: 268 posts, 5%. No decision stated: 1,254 posts, 22%.A third of the posts are decisionsabout data, and a fifth are games5,595 posts in six groups by what the decision isabout, shaded by what each post measuredNo measurementMeasured demoLabeled production, not confirmedProduction, confirmed by the auditDecisions about dataclassify, route, extract, rerank, judge, telemetryDecisions about data, no measurement: 1,135 postsDecisions about data, measured demo: 778 postsDecisions about data, labeled production, not confirmed: 19 postsDecisions about data, production, confirmed by the audit: 4 posts1,936 posts· 35% · 41% measuredDecisions inside games and control loopsgames, simulations, robotsDecisions inside games and control loops, no measurement: 791 postsDecisions inside games and control loops, measured demo: 280 posts1,071 posts· 19% · 26% measuredDecisions inside agent loopsgate a tool call, pick the next browser action, prune contextDecisions inside agent loops, no measurement: 394 postsDecisions inside agent loops, measured demo: 207 postsDecisions inside agent loops, labeled production, not confirmed: 1 postsDecisions inside agent loops, production, confirmed by the audit: 1 posts603 posts· 11% · 35% measuredDecisions about what a person just said or typedmoderation, voice, live chat, collaborationDecisions about what a person just said or typed, no measurement: 338 postsDecisions about what a person just said or typed, measured demo: 121 postsDecisions about what a person just said or typed, labeled production, not confirmed: 4 posts463 posts· 8% · 27% measuredDecisions about moneytrading botsDecisions about money, no measurement: 145 postsDecisions about money, measured demo: 122 postsDecisions about money, labeled production, not confirmed: 1 posts268 posts· 5% · 46% measuredNo decision statedmeta, benchmarks of the model, wrappersNo decision stated, no measurement: 1,032 postsNo decision stated, measured demo: 215 postsNo decision stated, labeled production, not confirmed: 4 postsNo decision stated, production, confirmed by the audit: 3 posts1,254 posts· 22% · 18% measuredSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
GroupPostsShareNo measurementMeasured demoLabeled production, not confirmedProduction, confirmedIn a live loopNeeds under 300 ms
Decisions about data (classify, route, extract, rerank, judge, telemetry)1,93635%1,1357781941%0%
Decisions inside games and control loops (games, simulations, robots)1,07119%7912800072%66%
Decisions inside agent loops (gate a tool call, pick the next browser action, prune context)60311%394207113%1%
Decisions about what a person just said or typed (moderation, voice, live chat, collaboration)4638%3381214045%18%
Decisions about money (trading bots)2685%1451221039%2%
No decision stated (meta, benchmarks of the model, wrappers)1,25422%1,032215431%0%
All posts5,595100%3,8351,723298

Read everything below knowing the week was mostly demos: two thirds of posts measured nothing, and a third measured a demo. Production numbers in posts are rare (8 of 5,595 held up on the audit), which fits the gateway picture: the volume sits with apps that don't post.

4 · Was any of it new?

The buzz is about using it, not about what it makes possible

Four in five posts compared Jev with nothing. That matters because Jev is doing jobs that already had tools. If it were solving something nobody could solve before, there'd be nothing to compare with. It isn't, so the missing baseline is the tell.

About four in five posts compared Jev with nothing, and about 1 in 30 with what it would replace5,595 posts. Audited: nothing 80% (77 to 82), the tools Jev would replace 3.7% (3 to 5). Nothing: 80.4%. Frontier LLM: 11.3%. Small LLM: 4.3%. Rules or regex: 2.5%. Classic classifier or ML: 1.1%. Vendor API: 0.6%. Audited: nothing 80% (77 to 82), frontier LLM 11.3% (10 to 13), small LLM 5.0% (4 to 7), the tools Jev would replace 3.7% (3 to 5).About four in five posts compared Jev with nothing, andabout 1 in 30 with what it would replace5,595 posts. Audited: nothing 80% (77 to 82), the tools Jev would replace 3.7% (3 to 5).0%25%50%75%100%Opus's labelsAudit estimate and 95% intervalNothingNothing: 80.4% of posts (4,498), Opus's labels80.4%4,498 postsNothing: audit 80.0% (95% interval 77.4% to 81.6%)audit 80% (77 to 82)Frontier LLMFrontier LLM: 11.3% of posts (630), Opus's labels11.3%630 postsFrontier LLM: audit 11.3% (95% interval 10.5% to 13.1%)audit 11.3% (10 to 13)Small LLMSmall LLM: 4.3% of posts (239), Opus's labels4.3%239 postsSmall LLM: audit 5.0% (95% interval 4.2% to 6.9%)audit 5.0% (4 to 7)Rules or regexRules or regex: 2.5% of posts (138), Opus's labels2.5%138 postsClassic classifier or MLClassic classifier or ML: 1.1% of posts (59), Opus's labels1.1%59 postsVendor APIVendor API: 0.6% of posts (31), Opus's labels0.6%31 posts4.1% combined (228 posts): the tools Jevwould replace. Audit: 3.7% (3 to 5)Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
About four in five posts compared Jev with nothing, and about 1 in 30 with what it would replace5,595 posts. Audited: nothing 80% (77 to 82), the tools Jev would replace 3.7% (3 to 5). Nothing: 80.4%. Frontier LLM: 11.3%. Small LLM: 4.3%. Rules or regex: 2.5%. Classic classifier or ML: 1.1%. Vendor API: 0.6%. Audited: nothing 80% (77 to 82), frontier LLM 11.3% (10 to 13), small LLM 5.0% (4 to 7), the tools Jev would replace 3.7% (3 to 5).About four in five posts comparedJev with nothing, and about 1 in 30with what it would replace5,595 posts. Audited: nothing 80% (77 to 82), thetools Jev would replace 3.7% (3 to 5).0%25%50%75%100%Opus's labelsAudit estimate and 95% intervalNothingNothing: 80.4% of posts (4,498), Opus's labels80.4%4,498 postsNothing: audit 80.0% (95% interval 77.4% to 81.6%)audit 80% (77 to 82)Frontier LLMFrontier LLM: 11.3% of posts (630), Opus's labels11.3%630 postsFrontier LLM: audit 11.3% (95% interval 10.5% to 13.1%)audit 11.3% (10 to 13)Small LLMSmall LLM: 4.3% of posts (239), Opus's labels4.3%239 postsSmall LLM: audit 5.0% (95% interval 4.2% to 6.9%)audit 5.0% (4 to 7)Rules or regexRules or regex: 2.5% of posts (138), Opus's labels2.5%138 postsClassic classifier or MLClassic classifier or ML: 1.1% of posts (59), Opus's labels1.1%59 postsVendor APIVendor API: 0.6% of posts (31), Opus's labels0.6%31 posts4.1% combined (228posts): the tools Jevwould replace. Audit:3.7% (3 to 5)Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
BaselinePostsShare of postsShare of viewsAudit estimate (95% interval)
Nothing4,49880.4%74.9%80% (77 to 82)
Frontier LLM63011.3%14.3%11.3% (10 to 13)
Small LLM2394.3%6.9%5.0% (4 to 7)
Rules or regex1382.5%2.2%
Classic classifier or ML591.1%0.5%
Vendor API310.6%1.3%
The tools Jev would replace (the last three)2284.1%3.7% (3 to 5)

The audit went further. It took the 91 posts where the first-pass labels found a measured decision inside a live system with a person waiting on it, the builds most likely to need both the speed and the price, and asked what a team would have used before Jev. An LLM for 64. Rules for 12. A vendor API for 11. A classic model for 4. None did something that was unavailable before. (85 distinct builds; 38 still meet the test on the auditor's own labels.)

Of the 91 builds most likely to show something new, none did something that was unavailable before91 builds (85 distinct), by what a team would have used before Jev. An LLM: 64 of 91. Rules or heuristics: 12 of 91. A vendor API: 11 of 91. A classic model: 4 of 91. Unavailable at any price: 0 of 91.Of the 91 builds most likely to show something new, nonedid something that was unavailable before91 builds (85 distinct), by what a team would have used before Jev.An LLMAn LLM: 64 of 91 posts6470%Rules or heuristicsRules or heuristics: 12 of 91 posts1213%A vendor APIA vendor API: 11 of 91 posts1112%A classic modelA classic model: 4 of 91 posts44%Unavailable at any priceUnavailable at any price: 0 of 91 posts0of 91 postsSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
Of the 91 builds most likely to show something new, none did something that was unavailable before91 builds (85 distinct), by what a team would have used before Jev. An LLM: 64 of 91. Rules or heuristics: 12 of 91. A vendor API: 11 of 91. A classic model: 4 of 91. Unavailable at any price: 0 of 91.Of the 91 builds most likely to showsomething new, none did somethingthat was unavailable before91 builds (85 distinct), by what a team would haveused before Jev.An LLMAn LLM: 64 of 91 posts6470%Rules or heuristicsRules or heuristics: 12 of 91 posts1213%A vendor APIA vendor API: 11 of 91 posts1112%A classic modelA classic model: 4 of 91 posts44%Unavailable at any priceUnavailable at any price: 0 of 91 posts0of 91 postsSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
What would have done the job beforePostsShare of the candidates
An LLM6470%
Rules or heuristics1213%
A vendor API1112%
A classic model44%
Unavailable at any price00%
Distinct builds among them85
Job: feed or comment filters2224%
Job: typing or form fill1314%
Job: voice commands1314%
Job: triage67%
Job: voice turn-taking67%
Job: guards55%
Job: speech scoring55%
Job: trading decisions55%
Job: classification33%
Job: computer use33%
Job: games or simulations33%
Job: alerts22%
Job: live control22%
Job: search22%
Job: personal tools11%

Faster and cheaper versions of jobs that had a tool. A real change to who can try them; one week in, not a change to what gets built.

5 · Announced, claimed, measured

TypeSafe's 40 to 400x is against frontier models, and holds there. Against the small models you'd actually use, the median is 7.3x.

The launch claim was against frontier models, and on the four published frontier comparisons it holds. The posts repeated it. Where someone put Jev against a small model on the same task and published the numbers (I found 12 such comparisons, including my own Pong runs) the gap shrank to a median of 7.3x on cost and 4.7x on latency (6x and 5.3x without my own runs), with Jev at or above the small model on accuracy for most bounded questions. Against frontier models the multiples are real: 97x to 580x on cost.

TypeSafe's 40 to 400x cheaper is against frontier models. Against comparable small models, the measured median is 7.3x.Cost and speed multiples: the launch claim, what the posts claimed, and what the published head-to-heads measured Launch claim 40 to 400x cheaper and 20 to 200x faster against frontier models. Posts: median 28x and 6x. Head-to-heads against small models: cost 2.8x to 58x, median 7.3x; latency 1.8x to 11x, median 4.7x.TypeSafe's 40 to 400x cheaper is against frontier models.Against comparable small models, the measured medianis 7.3x.Cost and speed multiples: the launch claim, what the posts claimed, and what the publishedhead-to-heads measuredCheaper by1x10x100x1000xTypeSafe, launch postTypeSafe, launch post: 40x to 400x40x to 400xagainst frontier LLMsThe posts' claims (126chips)The posts' claims (126 chips): 5.3x to 100x, median 28xmedian 28x5.3x to 100x, median 28xmiddle half and median; mostly against frontier LLMsMeasured, against a frontiermodel (4)Measured, against a frontier model (4): 97x to 580x97x to 580xMeasured, against a smallmodel (10)Measured, against a small model (10): 2.8x to 58x, median 7.3xmedian 7.3x2.8x to 58x, median 7.3xFaster by1x10x100x1000xTypeSafe, launch postTypeSafe, launch post: 20x to 200x20x to 200xagainst frontier LLMsThe posts' claims (297chips)The posts' claims (297 chips): 2x to 16x, median 6xmedian 6x2x to 16x, median 6xmiddle half and medianMeasured, against a frontiermodel (4)Measured, against a frontier model (4): 8.2x to 36x8.2x to 36xMeasured, against a smallmodel (11)Measured, against a small model (11): 1.8x to 11x, median 4.7xmedian 4.7x1.8x to 11x, median 4.7xThe launch claim and most chips compare with frontier models; the measured rows are every head-to-head withpublished numbers I could find, listed under the chart. Accuracy is in the table: Jev was at or above the small modelon most bounded questions.Source: the linked write-ups and the Jev Pong runs. Method: see the end of this page.
TypeSafe's 40 to 400x cheaper is against frontier models. Against comparable small models, the measured median is 7.3x.Cost and speed multiples: the launch claim, what the posts claimed, and what the published head-to-heads measured Launch claim 40 to 400x cheaper and 20 to 200x faster against frontier models. Posts: median 28x and 6x. Head-to-heads against small models: cost 2.8x to 58x, median 7.3x; latency 1.8x to 11x, median 4.7x.TypeSafe's 40 to 400x cheaper isagainst frontier models. Againstcomparable small models, themeasured median is 7.3x.Cost and speed multiples: the launch claim, whatthe posts claimed, and what the publishedhead-to-heads measuredCheaper by1x10x100x1000xTypeSafe, launchpostTypeSafe, launch post: 40x to 400x40x to 400xagainst frontier LLMsThe posts' claims(126 chips)The posts' claims (126 chips): 5.3x to 100x, median 28xmedian 28x5.3x to 100x, median 28xmiddle half and median; mostly against frontier LLMsMeasured, against afrontier model (4)Measured, against a frontier model (4): 97x to 580x97x to 580xMeasured, against asmall model (10)Measured, against a small model (10): 2.8x to 58x, median 7.3xmedian 7.3x2.8x to 58x, median 7.3xFaster by1x10x100x1000xTypeSafe, launchpostTypeSafe, launch post: 20x to 200x20x to 200xagainst frontier LLMsThe posts' claims(297 chips)The posts' claims (297 chips): 2x to 16x, median 6xmedian 6x2x to 16x, median 6xmiddle half and medianMeasured, against afrontier model (4)Measured, against a frontier model (4): 8.2x to 36x8.2x to 36xMeasured, against asmall model (11)Measured, against a small model (11): 1.8x to 11x, median 4.7xmedian 4.7x1.8x to 11x, median 4.7xThe launch claim and most chips compare with frontier models;the measured rows are every head-to-head with publishednumbers I could find, listed under the chart. Accuracy is in thetable: Jev was at or above the small model on most boundedquestions.Source: the linked write-ups and the Jev Pong runs. Method:see the end of this page.
Against small models the median gap is 7.3x on cost and 4.7x on latency; against frontier models it is 100x and moreEvery published head-to-head found: how many times cheaper and faster Jev was than the model it was compared with, on the same task 16 comparisons from 12 sources. Blue: against a small model. Grey: against a frontier model.Against small models the median gap is 7.3x on cost and4.7x on latency; against frontier models it is 100x andmoreEvery published head-to-head found: how many times cheaper and faster Jev was than themodel it was compared with, on the same taskcheaper by (against a small model)faster byagainst a frontier model1x10x100x1000xJev Pong (me) vs Ministral 3BJev Pong (me) vs Ministral 3B: 2.8x cheaperJev Pong (me) vs Ministral 3B: 1.8x fasterTaishi Morinaga vs Gemini 3.5 FlashTaishi Morinaga, Classmethod vs Gemini 3.5 Flash: 3.2x fasterCodeAlive vs gpt-oss-120bCodeAlive vs gpt-oss-120b: 4x cheaperCodeAlive vs gpt-oss-120b: 4.7x fasterJalil Laaraichi vs GPT-5.6 LunaJalil Laaraichi, OpenWork vs GPT-5.6 Luna: 5.2x cheaperanessbelbati vs Cohere Rerank 4 Proanessbelbati, rerank bench vs Cohere Rerank 4 Pro: 5.6x cheaperanessbelbati, rerank bench vs Cohere Rerank 4 Pro: 2x fasterEmil Lindfors vs DeepSeek V4.1 FlashEmil Lindfors vs DeepSeek V4.1 Flash: 6x cheaperEmil Lindfors vs DeepSeek V4.1 Flash: 8.4x fasterJon Reed vs Mistral Small 4Jon Reed, Near Here vs Mistral Small 4: 8.6x cheaperJon Reed, Near Here vs Mistral Small 4: 4.9x fasteranpicasso vs a small chat modelanpicasso, Hermes approvals vs a small chat model: 9.8x fasterJev Pong (me) vs GPT-5.4 NanoJev Pong (me) vs GPT-5.4 Nano: 10x cheaperJev Pong (me) vs GPT-5.4 Nano: 3.6x fasterDan Willoughby vs Claude Haiku 4.5Dan Willoughby, Sniff Test vs Claude Haiku 4.5: 33x cheaperDan Willoughby, Sniff Test vs Claude Haiku 4.5: 11x fasterJev Pong (me) vs Claude Haiku 4.5Jev Pong (me) vs Claude Haiku 4.5: 44x cheaperJev Pong (me) vs Claude Haiku 4.5: 3.6x fasterJon Reed vs Gemini 3.5 Flash-LiteJon Reed, Near Here vs Gemini 3.5 Flash-Lite: 58x cheaperJon Reed, Near Here vs Gemini 3.5 Flash-Lite: 5.8x fasterLangWatch vs Claude Opus 5LangWatch vs Claude Opus 5: 97x cheaperLangWatch vs Claude Opus 5: 8.2x fasterawlevin vs Claude Opus 5awlevin, computer use vs Claude Opus 5: 160x cheaperawlevin, computer use vs Claude Opus 5: 14x fasterDan Willoughby vs Claude Opus 5Dan Willoughby, Sniff Test vs Claude Opus 5: 239x cheaperDan Willoughby, Sniff Test vs Claude Opus 5: 36x fasterDan Shipper vs Claude Fable 5.1Dan Shipper, Every vs Claude Fable 5.1: 580x cheaperDan Shipper, Every vs Claude Fable 5.1: 25x fasterSmall: the vendor's small or cheap tier. Each multiple is the comparison model's figure divided by Jev's, as thesource published it; the table under the chart has the accuracy and the links.Source: the linked write-ups and the Jev Pong runs. Method: see the end of this page.
Against small models the median gap is 7.3x on cost and 4.7x on latency; against frontier models it is 100x and moreEvery published head-to-head found: how many times cheaper and faster Jev was than the model it was compared with, on the same task 16 comparisons from 12 sources. Blue: against a small model. Grey: against a frontier model.Against small models the median gapis 7.3x on cost and 4.7x on latency;against frontier models it is 100x andmoreEvery published head-to-head found: how manytimes cheaper and faster Jev was than the model itwas compared with, on the same taskcheaper by (against a small model)faster byagainst a frontier model1x10x100x1000xJev Pong (me) vsMinistral 3BJev Pong (me) vs Ministral 3B: 2.8x cheaperJev Pong (me) vs Ministral 3B: 1.8x fasterTaishi Morinaga vsGemini 3.5 FlashTaishi Morinaga, Classmethod vs Gemini 3.5 Flash: 3.2x fasterCodeAlive vsgpt-oss-120bCodeAlive vs gpt-oss-120b: 4x cheaperCodeAlive vs gpt-oss-120b: 4.7x fasterJalil Laaraichi vsGPT-5.6 LunaJalil Laaraichi, OpenWork vs GPT-5.6 Luna: 5.2x cheaperanessbelbati vs CohereRerank 4 Proanessbelbati, rerank bench vs Cohere Rerank 4 Pro: 5.6x cheaperanessbelbati, rerank bench vs Cohere Rerank 4 Pro: 2x fasterEmil Lindfors vsDeepSeek V4.1 FlashEmil Lindfors vs DeepSeek V4.1 Flash: 6x cheaperEmil Lindfors vs DeepSeek V4.1 Flash: 8.4x fasterJon Reed vs MistralSmall 4Jon Reed, Near Here vs Mistral Small 4: 8.6x cheaperJon Reed, Near Here vs Mistral Small 4: 4.9x fasteranpicasso vs a smallchat modelanpicasso, Hermes approvals vs a small chat model: 9.8x fasterJev Pong (me) vsGPT-5.4 NanoJev Pong (me) vs GPT-5.4 Nano: 10x cheaperJev Pong (me) vs GPT-5.4 Nano: 3.6x fasterDan Willoughby vs ClaudeHaiku 4.5Dan Willoughby, Sniff Test vs Claude Haiku 4.5: 33x cheaperDan Willoughby, Sniff Test vs Claude Haiku 4.5: 11x fasterJev Pong (me) vs ClaudeHaiku 4.5Jev Pong (me) vs Claude Haiku 4.5: 44x cheaperJev Pong (me) vs Claude Haiku 4.5: 3.6x fasterJon Reed vs Gemini 3.5Flash-LiteJon Reed, Near Here vs Gemini 3.5 Flash-Lite: 58x cheaperJon Reed, Near Here vs Gemini 3.5 Flash-Lite: 5.8x fasterLangWatch vs ClaudeOpus 5LangWatch vs Claude Opus 5: 97x cheaperLangWatch vs Claude Opus 5: 8.2x fasterawlevin vs Claude Opus 5awlevin, computer use vs Claude Opus 5: 160x cheaperawlevin, computer use vs Claude Opus 5: 14x fasterDan Willoughby vs ClaudeOpus 5Dan Willoughby, Sniff Test vs Claude Opus 5: 239x cheaperDan Willoughby, Sniff Test vs Claude Opus 5: 36x fasterDan Shipper vs ClaudeFable 5.1Dan Shipper, Every vs Claude Fable 5.1: 580x cheaperDan Shipper, Every vs Claude Fable 5.1: 25x fasterSmall: the vendor's small or cheap tier. Each multiple is thecomparison model's figure divided by Jev's, as the sourcepublished it; the table under the chart has the accuracy and thelinks.Source: the linked write-ups and the Jev Pong runs. Method:see the end of this page.
Show the numbers
WhoThe taskAgainstCheaper byFaster byAccuracySource
Jev Pong (me)the next paddle move, 30 statesMinistral 3B (small)2.8x1.8xJev 28 of 29, Ministral 26 of 30link
Jev Pong (me)the next paddle move, 30 statesGPT-5.4 Nano (small)10x3.6xJev 28 of 29, Nano 29 of 30link
Jev Pong (me)the next paddle move, 30 statesClaude Haiku 4.5 (small)44x3.6xJev 28 of 29, Haiku 30 of 30link
Jon Reed, Near Herereject unsuitable event listings, 50 casesMistral Small 4 (small)8.6x4.9xJev 96%, Mistral 84%link
Jon Reed, Near Herereject unsuitable event listings, 50 casesGemini 3.5 Flash-Lite (small)58x5.8xJev 96%, Gemini 86%link
Emil Lindfors11 typed questions over 24 hearing responsesDeepSeek V4.1 Flash (small)6x8.4xstance 20 of 24 each; substance Jev 19, DeepSeek 14link
Dan Willoughby, Sniff Testten boolean questions a paragraph, 54 clean paragraphsClaude Haiku 4.5 (small)33x11xfalse flags: Jev 1, Haiku 37link
Jalil Laaraichi, OpenWorkpage, ticket or ignore over 8,000 log linesGPT-5.6 Luna (small)5.2xnot publishedJev 100% recall and 0 false pages, Luna 96.2%link
CodeAliveblock or allow agent input, 58 messagesgpt-oss-120b (small)4x4.7xJev 9 of 9 hostile blocked and 0 of 49 real blocked; gpt-oss 8 or 9 of 9link
anessbelbati, rerank benchsearch reranking across 14 datasetsCohere Rerank 4 Pro (small)5.6x2xnDCG@10 0.692 against 0.691link
Taishi Morinaga, Classmethodfour-tier model routing, 40 callsGemini 3.5 Flash (small)not published3.2xJev 10 of 10 on every patternlink
anpicasso, Hermes approvalsapprove, deny or escalate 153 real commandsa small chat model (small)not published9.8x10 escalations to a human instead of 42link
LangWatchjudge 300 support conversations against human labelsClaude Opus 5 (frontier)97x8.2xJev 97% agreement, Opus 91%link
Dan Shipper, Everyfour writing checks over 12 passagesClaude Fable 5.1 (frontier)580x25xJev 6 of 7 planted defects, Fable 7 of 7link
awlevin, computer usethe next action from an OCR of the screenClaude Opus 5 (frontier)160x14xnot measuredlink
Dan Willoughby, Sniff Testten boolean questions a paragraph, 54 clean paragraphsClaude Opus 5 (frontier)239x36xfalse flags: Jev 1, Opus 0link

The cost uses people measured

JobWhoWhat they measuredSource
Evals and judgingLangWatch97% agreement with human labels over 300 support conversations; 10,000 conversations judged for $0.32 in 73 seconds, against $31 and ten minutes on Opus 5langwatch.ai
Evals and judgingMike Taylor and Dan Shipper, Every777 judgments in under 0.7 seconds; 0.35 s a passage against 8.8 s on Fable 5.1, about 580 times cheaperevery.to
ClassificationHassan, 1kpapers1,018 papers sorted into 24 topics for $0.08, 256 ms a paperX
ClassificationEmil Lindfors$0.22 per 1,000 documents; 0.32 s a documentlindfors.no
TriageJalil Laaraichi, OpenWorklog triage with 0 false pages; $0.062 against $0.32 per 3,000 logsdev.to
Security screeningGaurav Gosain96.5% on 662 blind prompt-injection and vulnerable-code cases; p50 325 msGitHub
Agent approvalsanpicasso, Hermes153 real commands: 10 escalations to a human instead of 42; 405 ms against 3,968 msGitHub
Computer useawlevinnext action from an OCR of the screen: $0.0002 a step against $0.032; 0.13 to 0.38 s against 5.2 sGitHub
Browser agentsGregor Zunic, Browser Usea real flight search 25% faster end to end; 101 protocol calls instead of 1,092GitHub
Context pruningtamarascore every tool call and result, drop the stale ones: 1M tokens to 86Kmadewithjev
Rerankinganessbelbatiparity with Cohere Rerank 4 Pro on nDCG at half the latencyGitHub

The honest number for a build-or-buy is single digits against a small model. Worth having, and it moves the sweet spot from ten times the decisions at a tenth of the bill to a few times the decisions at a fraction of it. The jobs are evals, triage, extraction, reranking and agent gating, and small-model prices have been falling roughly 10x a year on their own.

6 · Who needs it fast

Outside games, speed is a handful of demos, and they share one shape

Put what people built against how soon the decision has to land and there is one hot spot. Games. Everything else that needs an answer in under 300 ms fits in two small cells: typing and voice turn-taking. The volume sits with data decisions that can wait seconds, and with posts that never say what the decision is.

The need for a decision in under 300 ms has one hot spot: games5,595 posts by what the decision is about and how soon it is needed, on the labels. Shade is the number of posts; the audit estimated only the under-300 ms column and its games share. Decisions about data: under 300 ms 4, under a second 540, seconds or slower 1,239, unclear 153. Decisions inside games and control loops: under 300 ms 707, under a second 230, seconds or slower 74, unclear 60. Decisions inside agent loops: under 300 ms 5, under a second 39, seconds or slower 557, unclear 2. Decisions about what a person just said or typed: under 300 ms 84, under a second 255, seconds or slower 99, unclear 25. Decisions about money: under 300 ms 6, under a second 44, seconds or slower 194, unclear 24. No decision stated: under 300 ms 3, under a second 146, seconds or slower 71, unclear 1,034. Audited: 11.0% of posts need under 300 ms (9.5 to 12.5), 91% of them games (82 to 95).The need for a decision in under 300 ms has one hot spot:games5,595 posts by what the decision is about and how soon it is needed, on the labels. Shade isthe number of posts; the audit estimated only the under-300 ms column and its games share.under 2525+100+250+500+750+1,000+ postsUnclearOutlined: games, 91% of the under-300 ms posts on the auditUnder 300 msframe, feel, turnUnder asecondinteractionSeconds orslowertask, batchUnclearnot placedAll postsand shareDecisions about dataDecisions about data, under 300 ms: 4 posts, 0% of the group40%Decisions about data, under a second: 540 posts, 28% of the group54028%Search,classificationDecisions about data, seconds or slower: 1,239 posts, 64% of the group1,23964%ClassificationDecisions about data, unclear: 153 posts, 8% of the group1538%Classification1,93635%Decisions about data: 1,936 posts, 34.6% of all postsDecisions inside gamesand control loopsDecisions inside games and control loops, under 300 ms: 707 posts, 66% of the group70766%Games orsimulationsDecisions inside games and control loops, under a second: 230 posts, 21% of the group23021%Games orsimulationsDecisions inside games and control loops, seconds or slower: 74 posts, 7% of the group747%Games orsimulationsDecisions inside games and control loops, unclear: 60 posts, 6% of the group606%Games orsimulations1,07119%Decisions inside games and control loops: 1,071 posts, 19.1% of all postsDecisions inside agentloopsDecisions inside agent loops, under 300 ms: 5 posts, 1% of the group51%Decisions inside agent loops, under a second: 39 posts, 6% of the group396%Computer useDecisions inside agent loops, seconds or slower: 557 posts, 92% of the group55792%Tool-callgating,computer useDecisions inside agent loops, unclear: 2 posts, 0% of the group20%60311%Decisions inside agent loops: 603 posts, 10.8% of all postsDecisions about what aperson just said ortypedDecisions about what a person just said or typed, under 300 ms: 84 posts, 18% of the group8418%Typing or formfill, voiceturn-takingDecisions about what a person just said or typed, under a second: 255 posts, 55% of the group25555%Feed orcommentfiltersDecisions about what a person just said or typed, seconds or slower: 99 posts, 21% of the group9921%Feed orcommentfiltersDecisions about what a person just said or typed, unclear: 25 posts, 5% of the group255%Feed orcommentfilters4638%Decisions about what a person just said or typed: 463 posts, 8.3% of all postsDecisions about moneyDecisions about money, under 300 ms: 6 posts, 2% of the group62%Decisions about money, under a second: 44 posts, 16% of the group4416%TradingdecisionsDecisions about money, seconds or slower: 194 posts, 72% of the group19472%TradingdecisionsDecisions about money, unclear: 24 posts, 9% of the group249%2685%Decisions about money: 268 posts, 4.8% of all postsNo decision statedNo decision stated, under 300 ms: 3 posts, 0% of the group30%No decision stated, under a second: 146 posts, 12% of the group14612%Talk about themodelNo decision stated, seconds or slower: 71 posts, 6% of the group716%Talk about themodelNo decision stated, unclear: 1,034 posts, 82% of the group1,03482%Talk about themodel1,25422%No decision stated: 1,254 posts, 22.4% of all postsAll groups80914%1,25422%2,23440%1,29823%5,595Cells are the labelling model's calls, and how soon a decision is needed is the field the blind audit agrees with least(75% at seven tiers). The audit puts 11.0% of posts under 300 ms (9.5 to 12.5), not the labels' 14.5%, and 91% ofthose are games (82 to 95): the only two numbers here it confirms.Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
The need for a decision in under 300 ms has one hot spot: games5,595 posts by what the decision is about and how soon it is needed, on the labels. Shade is the number of posts; the audit estimated only the under-300 ms column and its games share. Decisions about data: under 300 ms 4, under a second 540, seconds or slower 1,239, unclear 153. Decisions inside games and control loops: under 300 ms 707, under a second 230, seconds or slower 74, unclear 60. Decisions inside agent loops: under 300 ms 5, under a second 39, seconds or slower 557, unclear 2. Decisions about what a person just said or typed: under 300 ms 84, under a second 255, seconds or slower 99, unclear 25. Decisions about money: under 300 ms 6, under a second 44, seconds or slower 194, unclear 24. No decision stated: under 300 ms 3, under a second 146, seconds or slower 71, unclear 1,034. Audited: 11.0% of posts need under 300 ms (9.5 to 12.5), 91% of them games (82 to 95).The need for a decision in under 300ms has one hot spot: games5,595 posts by what the decision is about and howsoon it is needed, on the labels. Shade is the numberof posts; the audit estimated only the under-300 mscolumn and its games share.under 2525+100+250+500+750+1,000+ postsUnclearOutlined: games, 91% of the under-300 ms posts on the auditUnder 300 msframe, feel, turnUnder 1 sinteractionSeconds+task, batchUnclearnot placedDecisions about data1,936 · 35%Decisions about data, under 300 ms: 4 posts, 0% of the group40%Decisions about data, under a second: 540 posts, 28% of the group54028%Decisions about data, seconds or slower: 1,239 posts, 64% of the group1,23964%Decisions about data, unclear: 153 posts, 8% of the group1538%Decisions inside games and control loops1,071 · 19%Decisions inside games and control loops, under 300 ms: 707 posts, 66% of the group70766%Decisions inside games and control loops, under a second: 230 posts, 21% of the group23021%Decisions inside games and control loops, seconds or slower: 74 posts, 7% of the group747%Decisions inside games and control loops, unclear: 60 posts, 6% of the group606%Decisions inside agent loops603 · 11%Decisions inside agent loops, under 300 ms: 5 posts, 1% of the group51%Decisions inside agent loops, under a second: 39 posts, 6% of the group396%Decisions inside agent loops, seconds or slower: 557 posts, 92% of the group55792%Decisions inside agent loops, unclear: 2 posts, 0% of the group20%Decisions about what a person just said or typed463 · 8%Decisions about what a person just said or typed, under 300 ms: 84 posts, 18% of the group8418%Decisions about what a person just said or typed, under a second: 255 posts, 55% of the group25555%Decisions about what a person just said or typed, seconds or slower: 99 posts, 21% of the group9921%Decisions about what a person just said or typed, unclear: 25 posts, 5% of the group255%Decisions about money268 · 5%Decisions about money, under 300 ms: 6 posts, 2% of the group62%Decisions about money, under a second: 44 posts, 16% of the group4416%Decisions about money, seconds or slower: 194 posts, 72% of the group19472%Decisions about money, unclear: 24 posts, 9% of the group249%No decision stated1,254 · 22%No decision stated, under 300 ms: 3 posts, 0% of the group30%No decision stated, under a second: 146 posts, 12% of the group14612%No decision stated, seconds or slower: 71 posts, 6% of the group716%No decision stated, unclear: 1,034 posts, 82% of the group1,03482%All groups 5,59580914%1,25422%2,23440%1,29823%Cells are the labelling model's calls, and how soon a decision isneeded is the field the blind audit agrees with least (75% at seventiers). The audit puts 11.0% of posts under 300 ms (9.5 to 12.5), notthe labels' 14.5%, and 91% of those are games (82 to 95): the onlytwo numbers here it confirms.Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
GroupUnder 300 msUnder a secondSeconds or slowerUnclearAll posts
Decisions about data4 (0%)540 (28%)1,239 (64%)153 (8%)1,936
Decisions inside games and control loops707 (66%)230 (21%)74 (7%)60 (6%)1,071
Decisions inside agent loops5 (1%)39 (6%)557 (92%)2 (0%)603
Decisions about what a person just said or typed84 (18%)255 (55%)99 (21%)25 (5%)463
Decisions about money6 (2%)44 (16%)194 (72%)24 (9%)268
No decision stated3 (0%)146 (12%)71 (6%)1,034 (82%)1,254
All groups809 (14.5%)1,254 (22.4%)2,234 (39.9%)1,298 (23.2%)5,595

The audit confirms the shape: a tenth of posts need a decision in under 300 ms, and nine in ten of those are games. Voice, live chat and collaboration, where a person is waiting, are 4% (3 to 7) of posts, and none is measured in production. The few live builds that exist mostly sit in one repo, Nader Dabit's jev-experiments, and they share a shape: several of your own questions on every event, inside the turn, from a model you didn't train. That's the one pattern nothing served before for anyone without an ML team.

About 11 percent of posts need a decision in under 300 ms, and about 9 in 10 of those are games5,595 posts. Audited: under 300 ms 11.0% (9.5 to 12.5), games 91% of those (82 to 95). Under 300 ms: 11.0% of posts (9.5 to 12.5); games 90.8% of those.About 11 percent of posts need a decision in under 300ms, and about 9 in 10 of those are games5,595 posts. Audited: under 300 ms 11.0% (9.5 to 12.5), games 91% of those (82 to 95).Under 300 ms: 11.0% of posts, 95% interval 9.5% to 12.5%Games under 300 ms: 10.0% of postsOther posts under 300 ms: 1.0% of postsEverything else: 89.0% of posts0%25%50%75%100%Share of the 5,595 postsGames under 300 msOther posts under 300 msEverything else95% intervalSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
About 11 percent of posts need a decision in under 300 ms, and about 9 in 10 of those are games5,595 posts. Audited: under 300 ms 11.0% (9.5 to 12.5), games 91% of those (82 to 95). Under 300 ms: 11.0% of posts (9.5 to 12.5); games 90.8% of those.About 11 percent of posts need adecision in under 300 ms, and about 9in 10 of those are games5,595 posts. Audited: under 300 ms 11.0% (9.5 to12.5), games 91% of those (82 to 95).Under 300 ms: 11.0% of posts, 95% interval 9.5% to 12.5%Games under 300 ms: 10.0% of postsOther posts under 300 ms: 1.0% of postsEverything else: 89.0% of posts0%25%50%75%100%Share of the 5,595 postsGames under 300 msOther posts under 300 msEverything else95% intervalSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
Audited estimateEstimate95% intervalFrom
Posts that need a decision in under 300 ms11.0%9.5% to 12.5%76 of 150 re-read posts still under 300 ms, plus 91 that only Opus puts there
The same, on the audit's samples alone10.3%8.2% to 14.3%neither model's word
Games, among those posts90.8%82.2% to 95.5%69 of 76
Outside games, about 8 percent of posts have a live loop, and none of them is in production376 posts flagged realtime outside games. Audited: a live loop 8.1% of posts (6.0 to 11.4). Labeled production: 2. Measured demo: 123. Demo, no numbers: 250. Proposal or commentary: 1.Outside games, about 8 percent of posts have a live loop,and none of them is in production376 posts flagged realtime outside games. Audited: a live loop 8.1% of posts (6.0 to 11.4).020406080100120Posts Opus flags realtimeLabeled production 2Measured demo 123Demo, no numbers 250Proposal or commentary 1Trading and marketsTrading and markets, labeled production: 1 postsTrading and markets, measured demo: 60 postsTrading and markets, demo, no numbers: 43 posts104Voice and turn-takingVoice and turn-taking, measured demo: 21 postsVoice and turn-taking, demo, no numbers: 48 posts69Live chat and streamsLive chat and streams, labeled production: 1 postsLive chat and streams, measured demo: 9 postsLive chat and streams, demo, no numbers: 54 posts64Collaboration and typingCollaboration and typing, measured demo: 11 postsCollaboration and typing, demo, no numbers: 51 posts62Browser and computer useBrowser and computer use, measured demo: 3 postsBrowser and computer use, demo, no numbers: 16 posts19Moderation and guardrailsModeration and guardrails, measured demo: 9 postsModeration and guardrails, demo, no numbers: 5 postsModeration and guardrails, proposal or commentary: 1 posts15Other or metaOther or meta, measured demo: 1 postsOther or meta, demo, no numbers: 14 posts15Classification and routingClassification and routing, measured demo: 4 postsClassification and routing, demo, no numbers: 5 posts9Data and telemetryData and telemetry, demo, no numbers: 7 posts7Search, rerank, extractionSearch, rerank, extraction, measured demo: 3 postsSearch, rerank, extraction, demo, no numbers: 3 posts6Evals and judgingEvals and judging, measured demo: 2 postsEvals and judging, demo, no numbers: 2 posts4Agent harness and gatingAgent harness and gating, demo, no numbers: 2 posts2Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
Outside games, about 8 percent of posts have a live loop, and none of them is in production376 posts flagged realtime outside games. Audited: a live loop 8.1% of posts (6.0 to 11.4). Labeled production: 2. Measured demo: 123. Demo, no numbers: 250. Proposal or commentary: 1.Outside games, about 8 percent ofposts have a live loop, and none ofthem is in production376 posts flagged realtime outside games. Audited:a live loop 8.1% of posts (6.0 to 11.4).020406080100120Posts Opus flags realtimeLabeled production 2Measured demo 123Demo, no numbers 250Proposal or commentary 1Trading and marketsTrading and markets, labeled production: 1 postsTrading and markets, measured demo: 60 postsTrading and markets, demo, no numbers: 43 posts104Voice and turn-takingVoice and turn-taking, measured demo: 21 postsVoice and turn-taking, demo, no numbers: 48 posts69Live chat and streamsLive chat and streams, labeled production: 1 postsLive chat and streams, measured demo: 9 postsLive chat and streams, demo, no numbers: 54 posts64Collaboration and typingCollaboration and typing, measured demo: 11 postsCollaboration and typing, demo, no numbers: 51 posts62Browser and computer useBrowser and computer use, measured demo: 3 postsBrowser and computer use, demo, no numbers: 16 posts19Moderation and guardrailsModeration and guardrails, measured demo: 9 postsModeration and guardrails, demo, no numbers: 5 postsModeration and guardrails, proposal or commentary: 1 posts15Other or metaOther or meta, measured demo: 1 postsOther or meta, demo, no numbers: 14 posts15Classification and routingClassification and routing, measured demo: 4 postsClassification and routing, demo, no numbers: 5 posts9Data and telemetryData and telemetry, demo, no numbers: 7 posts7Search, rerank, extractionSearch, rerank, extraction, measured demo: 3 postsSearch, rerank, extraction, demo, no numbers: 3 posts6Evals and judgingEvals and judging, measured demo: 2 postsEvals and judging, demo, no numbers: 2 posts4Agent harness and gatingAgent harness and gating, demo, no numbers: 2 posts2Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
FamilyLabeled productionMeasured demoDemo, no numbersProposal or commentaryFlagged realtime
Trading and markets160430104
Voice and turn-taking02148069
Live chat and streams1954064
Collaboration and typing01151062
Browser and computer use0316019
Moderation and guardrails095115
Other or meta0114015
Classification and routing04509
Data and telemetry00707
Search, rerank, extraction03306
Evals and judging02204
Agent harness and gating00202
All outside games21232501376

The live builds, by name

WhatWhoWhat they measuredSource
Voice turn-end on every partial transcriptNader Dabit108 ms round trip; the assistant replies 351 ms after the speaker stops instead of 1,000 msGitHub
Hold every chat message before it rendersNader Dabit45 messages a second held 124 ms each; 381 of 383 harmful messages caught, 0 of 958 clean ones blocked, on a labeled setGitHub
Judge a draft between keystrokesNader Dabit105 ms end to end, 85 judgments a second while typing; 4 of 4 risky drafts caught, 0 false blocksGitHub
Nine questions per customer message, eight chats at onceNader Dabitpanel refresh 93 msGitHub
Live minutes that flag a reversed decisionNader Dabit131 ms from end of utterance to screen; action items 96% recall, 92% precisionGitHub
A 300-message-a-second stream, every message judgedNader Dabit124 ms end to end, zero backlog, $37.86 an hourGitHub
Turn detection for voice agentsKwindla Hultman Kramer92.6% at 296 ms, against an LLM at 81.3% and 1,008 msX
Voice and gesture on a canvas, eight questions per partial transcriptgaborishka300 to 550 ms a decisionGitHub
Per-turn decisions as a LiveKit pluginsidxha proposal; 0 comments, 0 reactionsGitHub
The same shape, local, on a laptopmizorewww, laya-mlx7 to 14 ms a decision on an M3 Max; no cloud in the pathGitHub

Where speed and cost compound

The use cases people talked about, what they were done with before, and what a general decision model changes. Every one is a demo or a proposal.

Use caseHow it was done beforeWhat a decision model changesSpeed, cost, or bothEvidence so far
A voice control plane: speak now, is this a command, escalatesilence timers; a trained classifier per product; or an LLM at about a secondseveral product-specific questions on every partial transcript, about 300 ms, no trainingspeedtwo head-to-heads (92.6% at 296 ms against 81.3% at 1,008 ms); demos
A moderation policy per room, applied before fan-outkeyword filters; vendor taxonomies; an LLM per message at 0.5 to 3 s, so after the facta prose policy the room owner writes, several hazards per message, held about 120 msbotha synthetic set: 99.5% caught, 0 clean messages blocked
Who is this for? Gating an assistant in a groupwake words, @-mentions, or an LLM at about a seconda decision on every utterance at about 340 msspeeddemos, no accuracy numbers
A human-takeover trigger on every turna trained escalation classifier, or an LLM judge adding 1 to 3 sleft on for every turn, before the reply streamsboth10 escalations instead of 42, 405 ms against 3,968 ms; one measured demo
A live audience board, every message scoredsampled or batch sentiment; an LLM per message at $140 to $950 an hourevery message, categories redefined as you go, $37.86 an hour at 300 messages a secondcosta measured demo
Live minutes that flag a reversed decisiona summary after the meetingflagged in about 130 ms, while everyone is still in the roomcost, mostlyone repo, on replay

The latency story belongs to games and to local models: a local Laya answers in 7 to 14 ms on a laptop. The live human tier is where cost and speed compound into something new, and after a week it holds demos.

7 · Who else is coming

Twenty rivals in a week, most built by one person in days

The moat is smaller than the launch suggested. SemIf was up within a day. AutoJev was trained by agents in 20 hours on one H200 for $3,100. Reflex was one Shopify engineer over three days. JevK5, a week old, is second on JevBench. Each rival moves one or two knobs (15 of the other top-20 systems beat Jev on cost, 8 on speed, 1 on calibration, none on intelligence), and none matches Jev's mix yet. No big provider has announced a decision endpoint. And people built local copies partly because they couldn't get in.

Twenty rivals in a week, and the rate is still risingNew Hugging Face models named jev or laya a day, 17 to 23 Sep; posts in the feed showing a rival model, 16 to 23 Sep. The twenty are listed under the chart. 20 distinct models and endpoints plus 2 benchmarks by 22 Sep. Hugging Face: 1, 8, 17, 42, 57, 68, 77 a day.Twenty rivals in a week, and the rate is still risingNew Hugging Face models named jev or laya a day, 17 to 23 Sep; posts in the feed showing arival model, 16 to 23 Sep. The twenty are listed under the chart.Hugging Face: new models named jev or laya, per day02040608017 Sep: 1117 Sep18 Sep: 8818 Sep19 Sep: 171719 Sep20 Sep: 424220 Sep21 Sep: 575721 Sep22 Sep: 686822 Sep23 Sep: 777723 SepModels created that day whose name contains jev or laya.The feed: posts showing a model, port or endpoint that is not Jev, per day020406016 Sep: 6616 Sep17 Sep: 242417 Sep18 Sep: 343418 Sep19 Sep: 303019 Sep20 Sep: 474720 Sep21 Sep: 525221 Sep22 Sep: 434322 Sep23 Sep: 313123 SepAll Jev posting fell from 1,027 a day on 19 Sep to 336 on 23 Sep, so the rivals took a growing share of ashrinking conversation.Source: Hugging Face and jev.openchamber.dev, read 24 Sep 2026. Method: see the end of this page.
Twenty rivals in a week, and the rate is still risingNew Hugging Face models named jev or laya a day, 17 to 23 Sep; posts in the feed showing a rival model, 16 to 23 Sep. The twenty are listed under the chart. 20 distinct models and endpoints plus 2 benchmarks by 22 Sep. Hugging Face: 1, 8, 17, 42, 57, 68, 77 a day.Twenty rivals in a week, and the rateis still risingNew Hugging Face models named jev or laya a day,17 to 23 Sep; posts in the feed showing a rivalmodel, 16 to 23 Sep. The twenty are listed under thechart.Hugging Face: new models named jev or laya, per day02040608017 Sep: 1117 Sep18 Sep: 8818 Sep19 Sep: 171719 Sep20 Sep: 424220 Sep21 Sep: 575721 Sep22 Sep: 686822 Sep23 Sep: 777723 SepModels created that day whose name contains jev orlaya.The feed: posts showing a model, port or endpointthat is not Jev, per day020406016 Sep: 6616 Sep17 Sep: 242417 Sep18 Sep: 343418 Sep19 Sep: 303019 Sep20 Sep: 474720 Sep21 Sep: 525221 Sep22 Sep: 434322 Sep23 Sep: 313123 SepAll Jev posting fell from 1,027 a day on 19 Sep to 336on 23 Sep, so the rivals took a growing share of ashrinking conversation.Source: Hugging Face and jev.openchamber.dev, read 24 Sep2026. Method: see the end of this page.
Show the numbers
NameWhoFirst seenTypeCompetes onIndependent checkSource
SemIf (was OpenJev)TheoLeeCJ16 Sepopen-weights modelrun yourself, cost, accuracyJevBench #9 (was #2); 4,190 starslink
jevlikevinnylarouge16 Sepopen-weights modelrun yourself, cost, latencynone foundlink
decider-2bMapika16 Sepopen-weights model, 1.9Brun yourself, cost, latencyS1Bench: behind SimpleJev and Reflex among free optionslink
djev (DiffusionGemma as Jev)Maisa19 Sepopen-weights model and hosted endpointlatency, cost, generality, run yourselfJevBench #6link
openjev-sglangEric Zhang17 Sephosted endpointlatency, run yourselfnone foundlink
Qwen 3.8 27B on CerebrasShannon (@iamMrDuncan)17 Sephosted modellatency, accuracynone foundlink
KevJared Palmer17 Sepopen-weights models, 0.5B to 9Brun yourself, cost, accuracyJevBench #24 (4B), #45 (0.6B); 6,598 starslink
ReflexKshetrajna Raghavan (Shopify)17 Sepopen-weights model, in the browserrun yourself, costJevBench #5; cheaper than Jevlink
SimpleJevEugene Cheah (Featherless AI)18 Seplibrary and hosted endpoint, any model, visiongenerality, cost, run yourselfbest free option on S1Bench (19 Sep)link
LayaConvai Innovations18 Sepopen-weights model, 421Mlatency, cost, run yourselfJevBench #36; near chance zero-shot; a base to specializelink
localjevGitHub Next18 Seplocal port of Jev's APIrun yourselfnone foundlink
Verdict 2.0Hemant (@heman10x)18 Sepopen-weights model, 151M, in the browseraccuracy, cost, run yourselfJevBench #54link
DeepSeek V4.1 Flash behind a Jev-style endpointNick Khami19 Sephosted endpointgenerality, run yourselfnone foundlink
Distilled 4BTaro L. Saito19 Sepfine-tune recipelatency, accuracy, run yourselfnone foundlink
JeffLogan Markewich19 Sepself-hosted drop-inrun yourself, costJevBench #35link
laya-mlxmizorewww19 Seplocal port, Apple siliconlatency, run yourself7 to 14 ms on an M3 Max; 6,150 starslink
AutoJevDenis Yarats19 Sepopen-weights model, trained by agentsrun yourself, costnone found; 20 hours on one H200, $3,100 (the author)link
Winnow-12BEldan Ring20 Sepopen-weights vision modelgenerality, run yourselfJevBench #4link
HopperHopitAI21 Sepopen-weights modellatency, cost, calibration, run yourselfJevBench #3; better calibrated than Jevlink
JevK5allebee22 Sepopen-weights model, 4.2Blatency, cost, run yourselfJevBench #2, 62.0 against Jev 63.3link
JevBench (benchmark)Florian S, Benchmark Heaven19 Sepbenchmarkthe scoreboard77 systems ranked on 23 Seplink
S1Bench (benchmark)Cuth (@ItsCuthulhu)19 Sepbenchmarkthe scoreboard40 alternatives on 20 Seplink

What it took to get close: one person, a few days, open weights, and at most a few thousand dollars.

8 · Where it is

What the week adds up to

  • Jev is being called at scale, mostly by apps that don't say who they are, and the money is small.
  • Against comparable small models the gap is single digits, not 100x and 10x.
  • The uses that can be named are decisions about data: labeling, routing, judging, extraction. Possible before, cheaper now, and getting cheaper regardless.
  • Speed and cost compound in one place, decisions about a person's words in real time. It's 8% of posts, a few demos, and nothing measured in production.
  • None of the 91 builds most likely to show something new did something that was unavailable before.
  • Twenty rivals in a week, built cheaply, none matching the mix yet, and no big provider.

What to watch

  • 2 October. Jev stops being free on Vercel on 25 September. If its share of requests is still around a quarter a week later, the volume was use; if it halves, a lot of it was free-tier tinkering. I'll re-read the same export, OpenRouter's daily curve, and whether the five named apps are still there.
  • When TypeSafe reopens signups, and whether it publishes a usage number.
  • The first decision endpoint from a big provider.
  • The first production number in the live tier: a voice agent, a chat room, a takeover trigger, with a volume attached.

What I think it means is a separate piece. This page is the data.

Method

How this was measured

The posts come from OpenChamber's Jev feed, snapshot 23 September 18:47 UTC: 5,950 posts from 4,593 authors between 16 and 23 September, as OpenChamber selected them. Without 7 duplicates, the 347 posts that don't use Jev and one post the labeling model refused, 5,595 remain. The feed is what people chose to show, not a sample of usage: it starts six hours after launch, the top 1% of posts held half of all views, and the median post got 133 views.

I used AI models to do the sorting, and I want to be plain about that. Claude Opus 5.5 labeled every post against a written rubric: what Jev decides, whether the author measured anything, what they compared it with, whether the decision sits in a live loop, and whether it's in production. Then a second model, Grok, labeled 1,681 of those posts blind: every post in the rare groups the headlines rest on, and random samples of the rest. It produced its own estimates with 95% intervals, and wherever it checked a number this page uses its range, not the first model's count. It corrected several; the production count went from 37 to 8. How soon a decision is needed is the field it agrees with least, so the grid in section 6 shows the labels in three coarse bands with the audited 300 ms split as its headline, and its cells are the labels, not audited estimates. I hand-checked 13 posts myself, enough to catch problems, not enough to call it a human audit. A proper human sample is the check still missing. At Vercel AI Gateway list prices the labeling cost about $48.

The gateway numbers were read at source on 24 September: OpenRouter's model page and rankings, Vercel's open leaderboard export (CC BY 4.0), npm, pypistats and Discord. 24 September was a partial day and is left out everywhere. The head-to-heads are every comparison I could find where someone put Jev against another model on the same task and published cost, latency or accuracy; each is the author's own figure, unreproduced. The rivals were found from the feed, Hugging Face, GitHub and JevBench; ranks move daily and are dated. Gateway requests are requests, not decisions, and can't separate production from testing. Vercel publishes shares, never counts.

The labels, the tables behind every chart, the rubric and the code are at github.com/mattheworiordan/jev-landscape, without the text of any post: code under MIT, data and method under CC BY 4.0. The posts belong to their authors. The full technical page has every chart from the first edition, including the ones this page leaves out.

Disclosure. I'm CEO of Ably, a realtime infrastructure company. I looked at Jev because it sits in the low-latency part of the stack I work on. Read the numbers with that in mind. I'm on LinkedIn if you want to argue with any of it.

Show the audit tables

The tables refer to the figure numbers of the full technical page.

The audit

A piece about unmeasured claims should show its own measurements being checked. An independent reviewer, Grok, labeled 1,681 posts against the same rubric without seeing the model's labels: all 422 posts in the rare groups the headlines rest on, and seeded random samples of 100 to 300 posts from the rest. Each sample is scaled up to its group with a 95% Wilson interval. The review is published with every label and the scripts that turn them into the numbers on this page. The audit sampled from the Sonnet 5 labels. Because the charts use Opus's, I reran its estimators with Opus as the model being checked (the rerun): the audit's labels stay the reference, shares are of Opus's 5,595 posts, and where the audit didn't sample, Opus's labels fill in, with the count stated.

It tested ten of the claims I'd published: 4 hold, 3 hold with a correction and 3 fail. The failures are why this page no longer sorts posts into buckets. "Cost-only" (23.3% of posts) was the bin every other measured post fell into, and its description fits about 12.5% (9.8 to 16.6). And not one of the 91 "materially different" builds did something that was unavailable before. The other two fails, the share of posts that compare Jev with nothing and the share that are meta, rest on the audit taking the labels as right wherever it didn't sample. On its own random samples, the Sonnet 5 labels' 81.4% and 21.1% sit inside the intervals (77 to 85, and 16 to 25). That's why the audited numbers on this page are ranges.

What was checked

GroupHowPosts in the groupAudited
Measured production (Sonnet label)every post1515
The 91 candidate buildsevery post9191
Voice, live chat or collaborationevery post156156
Classic ML, rules or vendor baselineevery post201201
No comparison namedrandom sample4,647300
Demo, no numbersrandom sample3,681200
Measured demorandom sample1,928200
Not a realtime family, not gamesrandom sample4,475200
Realtime flag falserandom sample3,971200
Realtime flag true, not gamesrandom sample720100
Other or metarandom sample1,203150
Under 300 ms (frame, feel or turn)random sample1,064150
Unique posts audited1,681 (422 in the census groups)

How often the audit agrees with the model

The share of audited posts where the audit's label matches the Claude Opus 5.5 label, field by field. The groups aren't a random sample of the week, so the last row is agreement on the audited posts, not on every post.

GroupPostsFamilyTierEvidenceBaselineRealtimeProduction
Measured production (Sonnet label)1587%73%80%93%100%93%
The 91 candidate builds9184%75%89%92%88%99%
Voice, live chat or collaboration15679%75%97%97%93%99%
Classic ML, rules or vendor baseline20182%69%95%79%97%99%
No comparison named30082%75%93%96%96%98%
Demo, no numbers20082%71%92%98%92%98%
Measured demo20084%73%94%93%96%100%
Not a realtime family, not games20082%80%91%96%97%98%
Realtime flag false20076%80%95%92%98%98%
Realtime flag true, not games10083%78%95%94%90%96%
Other or meta15079%83%89%95%97%99%
Under 300 ms (frame, feel or turn)15093%72%97%97%89%99%
All audited cards1,68182%75%93%94%95%98%

The Claude Sonnet 5 labels the audit sampled from agree less on every field:

FieldSonnet 5 (first labels)Opus 5.5 (this page)
Family71% (κ 0.67)82% (κ 0.79)
Tier62% (κ 0.52)75% (κ 0.69)
Evidence89% (κ 0.79)93% (κ 0.86)
Baseline89% (κ 0.73)94% (κ 0.83)
Realtime86% (κ 0.67)95% (κ 0.86)
Production98% (κ 0.52)98% (κ 0.64)

The ten claims

As first published, with the audit's corrected value and its 95% interval.

#ClaimPublishedAudit on SonnetVerdictAudit on OpusVerdict on Opus
1Comparison baselines81.5% compare Jev with nothing, 11.9% with a frontier LLM, 3.0% with a small LLM, 3.6% with the thing Jev would replace78.2% nothing, 13.5% frontier LLM, 4.8% small LLM, 3.5% the tools Jev would replace; 75.7–79.8%; 12.8–15.3%; 3.9–6.6%; 2.8–5.1%Fails80.0% (77.4–81.6%); 11.3% (10.5–13.1%); 5.0% (4.2–6.9%); 3.7% (3.0–5.4%)Holds
2Production with numbersat most 1.6%; 0.9% on a re-read; 0.3% on the strict label0.14% (8 of 15 strict cards hold); 0.14–2.0%Holds with correction0.14% (0.14–2.01%)Holds with correction
3The bucket schemeabout 38% demos with no numbers; about 12% hype; about 23% cost-only; 1.6% (90 posts) materially different, 87 outside gamesthe demo and hype rules rest on the framing field, since withdrawn; 12.5% of posts fit the cost-only description, not 23.3%; 38 of the 91 still meet the material test (0.67%; 2.0% with misses), and none did something unavailable before; 9.8–16.6% cost-only; 1.2–4.6% materialFailsno measurement 67.3% (63.5–70.5%); measured demo 32.5% (29.4–36.3%); demos, no numbers (chip rule) 46.6% (41.8–51.6%); claim, no number (chip rule) 2.5% (1.3–5.4%); cost-only as described 12.5% (9.7–16.6%); material test 2.0% (1.2–4.6%)The ladder shares hold
4Voice, live chat and collaboration2.8% of posts, none in production3.8% of posts; 0 in production; 0 claiming production; 2.7–6.3%Holds with correction4.0% (2.9–6.6%); 0 in measured productionHolds
5A live loopabout a fifth of posts; 5 to 10 percent outside games7.8% outside games; about a fifth with games (21.7%); 5.8–11.1% outside gamesHoldsoutside games 8.1% (6.0–11.4%); all 22.4%Holds
6Decisions under 300 msof posts that need a decision in under 300 ms, about 91% are games90.8% are games; the set is 9.4% of posts, not 19%; 82.2–95.5% games; 8.0–10.9% of postsHoldsset 11.0% (9.5–12.5%); games 90.8% (82.2–95.5%)Holds; the set's size does not
7Posts that never say what Jev decidesa fifth never say what Jev decides, or are benchmarks or wrappers; memes and hot takes about 1% each14.9% stay meta; 12.4% vague, a benchmark or a wrapper; memes 0.28%, hot takes 0.42%; 13.3–16.3% still metaFailsmeta 20.9% (19.4–22.2%); memes 0.29%; hot takes 0.44%Holds with correction: the fifth holds on the audit-only estimate; memes and hot takes are under half a percent each
8Attentionhalf of all views sit on 1% of poststhe top 1% (58 posts) hold 53.4% of views; 49.6% without the 9 low like-rate posts; 47.5% of likes; a census, no samplingHolds with correctiontop 1% (56 posts) hold 53.3% of views; 49.8% without the low like-rate postsHolds with correction
9Jev as a classifierJev agreed with the model on 77% of cards, and was right 27 of 28 times at 0.99 or more76.7% (4,561 of 5,943); 27 of 28; 98.1% on the audit's 464 cards at 0.99 or more; a census of Jev's callsHoldsJev = Opus 78.4%; 97.0% at 0.99Holds
10The claimed multiplesthe median claim is about 28× cheaper and 6× faster28× cheaper (126 chips), 6× faster (299 chips); 34× and 8.5× without the 1× chips; a census of the chipsHolds28x cheaper (126 chips), 6x faster (297 chips)Holds

The numbers this page uses

Each audited estimate with its 95% interval, as a share of Opus's 5,595 posts. "Samples only" uses the audit's random samples where it didn't look, instead of Opus's labels. The last column is the audit's own estimate on the Sonnet 5 labels.

Posts thatAuditedSamples onlyOpus labelsAudit on Sonnet 5
Compare Jev with nothing80.0% (77.4 to 81.6)80.7% (76.8 to 84.8)80.4%78.2% (75.7 to 79.8)
Compare it with a frontier LLM11.3% (10.5 to 13.1)10.3% (6.9 to 14.2)11.3%13.5% (12.8 to 15.3)
Compare it with a small LLM5.0% (4.2 to 6.9)5.3% (2.9 to 9.9)4.3%4.8% (3.9 to 6.6)
Compare it with the tools it would replace3.7% (3.0 to 5.4)3.7% (2.9 to 7.3)4.1%3.5% (2.8 to 5.1)
Measured nothing67.3% (63.5 to 70.5)68.5%67.2% (63.4 to 70.3)
Demo with no numbers46.6% (41.8 to 51.6)47.9%46.5% (41.6 to 51.4)
A cost, speed or accuracy claim with no number2.5% (1.3 to 5.4)2.1%2.4% (1.2 to 5.3)
Measured demo32.5% (29.4 to 36.3)30.8%32.7% (29.6 to 36.5)
Measured production0.14% (0.14 to 2.01)0.66%0.14% (0.14 to 1.99)
Voice, live chat or collaboration4.0% (2.9 to 6.6)3.9% (2.7 to 7.9)3.7%3.8% (2.7 to 6.3)
A live loop outside games8.1% (6.0 to 11.4)8.3% (5.9 to 13.3)6.7%7.8% (5.8 to 11.1)
Need a decision in under 300 ms11.0% (9.5 to 12.5)10.3% (8.2 to 14.3)14.5%9.4% (8.0 to 10.9)
Games, of those under 300 ms90.8% (82.2 to 95.5)88.1%87.4%90.8% (82.2 to 95.5)
Meta (no use case, a benchmark or a wrapper)20.9% (19.4 to 22.2)20.7% (16.9 to 25.9)22.4%14.9% (13.3 to 16.3)

What the audit couldn't check

The audit drew its groups on the Sonnet 5 labels, so for Opus's labels some groups are thin: it read only part of the posts Opus puts in voice, live chat and collaboration, for example. The 847 posts the Sonnet 5 labels say compare Jev with a frontier or small LLM weren't sampled, so those estimates take Opus's word there. The posts from those groups the audit happened to label for other reasons agree 115 of 151 and 27 of 40 with the Sonnet 5 labels, and that slice isn't a random sample. The audit re-read the 15 posts the Sonnet 5 labels called production; its first-pass labels call three more posts production, which its count leaves out (Opus calls all three production too), and the 2% upper end allows for posts like those. The first review's re-read of the first pass's 95 production posts wasn't repeated. Games were left out of the check for missed voice, live chat and collaboration posts. The 78 proposal and commentary posts stayed on the model's label. The feed cuts post text at 400 characters and a claim chip can invent a number: 48 of the 1,023 posts the audit calls demos with no numbers still carry a numeric chip. And the auditor is another AI model, not a person. I labeled 13 posts myself, enough to drop two labels but not enough to check the rest, so human labels are still the missing check.

How the posts were labeled

  • Data. OpenChamber's Jev feed, snapshot 23 September 18:47 UTC: 5,950 posts from 4,593 authors, which OpenChamber selected with its own filter for what counts as a build. That includes 6 posts an earlier snapshot held and the feed later dropped. Post times decoded from the X ids run from 16 September 00:28 to 23 September 18:44 UTC. Views are X impressions and likes are X likes, both as the feed recorded them.
  • Open data. The labels from every model and the audit, the tables behind every chart, the rubric and the code are at github.com/mattheworiordan/jev-landscape. The feed's post text isn't republished there; the repo says how to fetch it.
  • Rubric. The v2 rubric has 15 families (the fourteen on the charts, and one for posts that don't use Jev), 7 latency tiers, 5 evidence levels, 5 framings, 6 baselines, and two flags: realtime infrastructure and a production claim. The framings and the seven tiers were labeled but aren't charted; they are in the labels file. "Measured" needs a number from the author's own run; TypeSafe's launch numbers quoted as Jev's general speed or price don't count. "Unclear" is allowed and preferred to a guess.
  • Models. Claude Opus 5.5, with adaptive thinking, labeled every post through Vercel AI Gateway, in batches of 40, with the v2 rubric as a cached system prompt (data/classified-opus.jsonl). Its safety filter refused one post, which is left out. Claude Sonnet 5 made the first two passes, and its v2 labels are the ones the audit sampled from. The Gateway ignores temperature for both models, so both runs were sampled at its default and aren't deterministic: on the same 120 cards, two Sonnet runs matched on family for 102 of 120. Jev classified the family of every post again as a second classifier, one call per post.
  • Noise sub-types. A separate Claude Sonnet 5 pass sorted the noise posts (Other or meta, or commentary in another family) into sub-types, and the hot takes by stance, against report/rubric-noise.md.
  • Cost. At Gateway list prices the first Sonnet pass cost $7.71 and the second $8.76 ($8.39 for the full run and the refreshes, $0.37 for two 120-card pilots). The Claude Opus 5.5 run cost $29.26. Sorting the noise into sub-types (figure 1b of the technical page) cost $1.47, and Jev's pass $0.33.
  • Base. 7 duplicate posts were merged, and the 347 posts that don't use Jev (5.8%: local clones, distillations, Jev-compatible APIs over other models, builds on other models) are left out of every use-case chart, as is the post Opus refused. That leaves 5,595 posts.
  • Agreement. The audit is the main check: Grok labeled 1,681 posts blind, and agrees with Opus on family for 82% of them, on evidence for 93% and on tier for 75%. An earlier check, 120 random posts labeled blind by Claude Opus 5.5 in a separate run, is in the report. I labeled 13 posts myself as a calibration. The readers that checked every number are AI models; a proper human sample is the check still missing.

The buckets are gone

The first versions of this page put every post in one of six buckets: noise, hype, demo, cost-only, fast loop and material. The rules were finished after the data arrived, and the audit showed they didn't hold (claim 3), so the buckets are withdrawn. The measurement ladder in figure 1 of the technical page and the substance test in figure 1c of the technical page replace them. The bucket tables are kept for the record in the repo (report/v1/buckets-on-v2-labels.md), and so is the Sonnet version of this analysis (report/v2/).

What changed from the first pass

A critical review of the first pass found that "hype" rested on Sonnet calling any build "capability", that the accuracy and latency chip medians mixed Jev's numbers with baselines, and that "5,782 people" was 5,782 posts from 4,469 authors. The v2 rubric moved measured production from 1.6% to 0.2% of posts and production claims from 4.7% to 1.8%. The realtime flag went the other way (25.8% to 29.8%) because v2 flags nearly every game; figure 7 of the technical page corrects for that with the audit.

Data quality and limits

  • Post text in the feed is capped at 400 characters, and 983 posts (17%) are cut. The claim chips come from the full post, so some numbers are visible only as chips.
  • The most-viewed post has a like rate of 0.06%, and 9 posts with 100,000 or more views and a like rate under 0.2% hold 12.4% of views. Figure 2 gives the numbers without them.
  • Every label is one model's reading of a short post, not a check of what was built. Claims on the cards are the authors' own and aren't reproduced here.
  • Every table behind these charts, the rubric and the scripts are in the repo; this page and its charts are generated from those CSVs by report/site/scripts/charts.mjs.