Data · 5,950 posts · 16 to 23 September 2026

A week of Jev, sorted

What 5,950 posts from 4,593 authors built with Jev in its first week, where the attention went, how much of it was measured, and whether any of it was new.

Every post that uses Jev, on the measurement ladder (figure 1):

  • no measurement3,835
  • measured demo1,723
  • labeled production, not confirmed29
  • production, confirmed8

Jev, TypeSafe AI's typed-decision model, launched on 15 September, the first of a new category that everyone got excited about at once. It doesn't write prose: it answers typed questions about a state (pick one of these, score this, yes or no), and it answers fast.

I went looking for where it applies to realtime, the part of the stack I spend my days on, so I sorted every post in OpenChamber's Jev feed from that first week. Wherever a blind audit of 1,681 posts checked a number, I give its range, not a model's count.

The short version

  • None of the 91 builds most likely to show something new did something that was unavailable before: a team would have used an LLM for 64 of them (figure 1c).
  • About four in five posts compare Jev with nothing, 80% (77 to 82), and about 1 in 30 with the tools it would replace, 3.7% (3 to 5).
  • Eight of 5,595 posts measured Jev in production, 0.14% (at most 2%), and none of them is a realtime build.
  • When Jev was surest it agreed with Opus 98.0% of the time, and it labeled every post for $0.33 against $29.26 for Opus.

Figure 1

The week in one picture

Two thirds of the posts measured nothing, 67% (64 to 70) on the audit, and about a third reported a number from the author's own run, 33% (29 to 36). Production is only the 8 orange squares, the posts that held when the audit re-read them: 0.14% of posts, and at most 2%.

About a third of the week's posts measured a demo, and 8 of 5,595 measured production5,595 posts. Audited: measured demo 33% (29 to 36), production 0.14% (at most 2%). No measurement: 3,835 posts. Measured demo: 1,723 posts. Labeled production, not confirmed: 29 posts. Production, confirmed: 8 posts.About a third of the week's posts measured a demo, and 8of 5,595 measured production5,595 posts. Audited: measured demo 33% (29 to 36), production 0.14% (at most 2%).No measurement 3,835Measured demo 1,723Labeled production, not confirmed 29Production, confirmed 8Other or meta1,254 posts · 18% measuredOther or meta, no measurement: 1,032 postsOther or meta, measured demo: 215 postsOther or meta, labeled production, not confirmed: 4 postsOther or meta, production, confirmed: 3 postsGames and control loops1,071 posts · 26% measuredGames and control loops, no measurement: 791 postsGames and control loops, measured demo: 280 postsClassification and routing781 posts · 40% measuredClassification and routing, no measurement: 467 postsClassification and routing, measured demo: 304 postsClassification and routing, labeled production, not confirmed: 8 postsClassification and routing, production, confirmed: 2 postsSearch, rerank, extraction518 posts · 39% measuredSearch, rerank, extraction, no measurement: 315 postsSearch, rerank, extraction, measured demo: 195 postsSearch, rerank, extraction, labeled production, not confirmed: 6 postsSearch, rerank, extraction, production, confirmed: 2 postsEvals and judging366 posts · 30% measuredEvals and judging, no measurement: 257 postsEvals and judging, measured demo: 106 postsEvals and judging, labeled production, not confirmed: 3 postsData and telemetry271 posts · 65% measuredData and telemetry, no measurement: 96 postsData and telemetry, measured demo: 173 postsData and telemetry, labeled production, not confirmed: 2 postsTrading and markets268 posts · 46% measuredTrading and markets, no measurement: 145 postsTrading and markets, measured demo: 122 postsTrading and markets, labeled production, not confirmed: 1 postsAgent harness and gating262 posts · 30% measuredAgent harness and gating, no measurement: 183 postsAgent harness and gating, measured demo: 77 postsAgent harness and gating, labeled production, not confirmed: 1 postsAgent harness and gating, production, confirmed: 1 postsModeration and guardrails257 posts · 32% measuredModeration and guardrails, no measurement: 176 postsModeration and guardrails, measured demo: 78 postsModeration and guardrails, labeled production, not confirmed: 3 postsBrowser and computer use254 posts · 37% measuredBrowser and computer use, no measurement: 159 postsBrowser and computer use, measured demo: 95 postsCompaction and context87 posts · 40% measuredCompaction and context, no measurement: 52 postsCompaction and context, measured demo: 35 postsVoice and turn-taking72 posts · 29% measuredVoice and turn-taking, no measurement: 51 postsVoice and turn-taking, measured demo: 21 postsLive chat and streams68 posts · 18% measuredLive chat and streams, no measurement: 56 postsLive chat and streams, measured demo: 11 postsLive chat and streams, labeled production, not confirmed: 1 postsCollaboration and typing66 posts · 17% measuredCollaboration and typing, no measurement: 55 postsCollaboration and typing, measured demo: 11 postsSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
About a third of the week's posts measured a demo, and 8 of 5,595 measured production5,595 posts. Audited: measured demo 33% (29 to 36), production 0.14% (at most 2%). No measurement: 3,835 posts. Measured demo: 1,723 posts. Labeled production, not confirmed: 29 posts. Production, confirmed: 8 posts.About a third of the week's postsmeasured a demo, and 8 of 5,595measured production5,595 posts. Audited: measured demo 33% (29 to36), production 0.14% (at most 2%).No measurement 3,835Measured demo 1,723Labeled production, not confirmed 29Production, confirmed 8Other or meta1,254 posts · 18% measuredOther or meta, no measurement: 1,032 postsOther or meta, measured demo: 215 postsOther or meta, labeled production, not confirmed: 4 postsOther or meta, production, confirmed: 3 postsGames and control loops1,071 posts · 26% measuredGames and control loops, no measurement: 791 postsGames and control loops, measured demo: 280 postsClassification and routing781 posts · 40% measuredClassification and routing, no measurement: 467 postsClassification and routing, measured demo: 304 postsClassification and routing, labeled production, not confirmed: 8 postsClassification and routing, production, confirmed: 2 postsSearch, rerank, extraction518 posts · 39% measuredSearch, rerank, extraction, no measurement: 315 postsSearch, rerank, extraction, measured demo: 195 postsSearch, rerank, extraction, labeled production, not confirmed: 6 postsSearch, rerank, extraction, production, confirmed: 2 postsEvals and judging366 posts · 30% measuredEvals and judging, no measurement: 257 postsEvals and judging, measured demo: 106 postsEvals and judging, labeled production, not confirmed: 3 postsData and telemetry271 posts · 65% measuredData and telemetry, no measurement: 96 postsData and telemetry, measured demo: 173 postsData and telemetry, labeled production, not confirmed: 2 postsTrading and markets268 posts · 46% measuredTrading and markets, no measurement: 145 postsTrading and markets, measured demo: 122 postsTrading and markets, labeled production, not confirmed: 1 postsAgent harness and gating262 posts · 30% measuredAgent harness and gating, no measurement: 183 postsAgent harness and gating, measured demo: 77 postsAgent harness and gating, labeled production, not confirmed: 1 postsAgent harness and gating, production, confirmed: 1 postsModeration and guardrails257 posts · 32% measuredModeration and guardrails, no measurement: 176 postsModeration and guardrails, measured demo: 78 postsModeration and guardrails, labeled production, not confirmed: 3 postsBrowser and computer use254 posts · 37% measuredBrowser and computer use, no measurement: 159 postsBrowser and computer use, measured demo: 95 postsCompaction and context87 posts · 40% measuredCompaction and context, no measurement: 52 postsCompaction and context, measured demo: 35 postsVoice and turn-taking72 posts · 29% measuredVoice and turn-taking, no measurement: 51 postsVoice and turn-taking, measured demo: 21 postsLive chat and streams68 posts · 18% measuredLive chat and streams, no measurement: 56 postsLive chat and streams, measured demo: 11 postsLive chat and streams, labeled production, not confirmed: 1 postsCollaboration and typing66 posts · 17% measuredCollaboration and typing, no measurement: 55 postsCollaboration and typing, measured demo: 11 postsSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
FamilyNo measurementMeasured demoLabeled production, not confirmedProduction, confirmedPosts
Other or meta1,032215431,254
Games and control loops791280001,071
Classification and routing46730482781
Search, rerank, extraction31519562518
Evals and judging25710630366
Data and telemetry9617320271
Trading and markets14512210268
Agent harness and gating1837711262
Moderation and guardrails1767830257
Browser and computer use1599500254
Compaction and context52350087
Voice and turn-taking51210072
Live chat and streams56111068
Collaboration and typing55110066
All posts3,835 (68.5%)1,723 (30.8%)29 (0.5%)8 (0.1%)5,595

Figure 1b

What the noise is made of

About a fifth of posts state no use case, benchmark the model itself or wrap it: 21% (19 to 22) on the audit. Of the 1,265 noise posts, 89% are unrelated or unclear posts, benchmarks of the model and tooling or wrappers, and memes and hot takes are under half a percent of all posts each.

About a fifth of posts are meta, and memes and hot takes are under half a percent each1,265 noise posts by sub-type. Audited: meta 21% (19 to 22) of all posts. Unrelated or unclear: 523 posts (41.3%), 33.1% of the noise's views. Benchmarks of the model: 360 posts (28.5%), 33.0% of the noise's views. Tooling or wrappers: 245 posts (19.4%), 19.6% of the noise's views. Explainers or tutorials: 59 posts (4.7%), 1.7% of the noise's views. News or reposts: 45 posts (3.6%), 10.6% of the noise's views. Hot takes: 18 posts (1.4%), 1.9% of the noise's views. Memes or jokes: 15 posts (1.2%), 0.1% of the noise's views.About a fifth of posts are meta, and memes and hot takesare under half a percent each1,265 noise posts by sub-type. Audited: meta 21% (19 to 22) of all posts.0%25%50%Share of the noiseShare of the noise postsShare of the noise viewsUnrelated or unclearUnrelated or unclear: 523 posts, 41.3% of the noiseUnrelated or unclear: 2,176,307 views, 33.1% of the noise's views41.3%523 posts33.1% of viewsBenchmarks of the modelBenchmarks of the model: 360 posts, 28.5% of the noiseBenchmarks of the model: 2,169,703 views, 33.0% of the noise's views28.5%360 posts33.0% of viewsTooling or wrappersTooling or wrappers: 245 posts, 19.4% of the noiseTooling or wrappers: 1,291,340 views, 19.6% of the noise's views19.4%245 posts19.6% of viewsExplainers or tutorialsExplainers or tutorials: 59 posts, 4.7% of the noiseExplainers or tutorials: 113,146 views, 1.7% of the noise's views4.7%59 posts1.7% of viewsNews or repostsNews or reposts: 45 posts, 3.6% of the noiseNews or reposts: 695,749 views, 10.6% of the noise's views3.6%45 posts10.6% of viewsHot takesHot takes: 18 posts, 1.4% of the noiseHot takes: 123,574 views, 1.9% of the noise's views1.4%18 posts1.9% of viewsMemes or jokesMemes or jokes: 15 posts, 1.2% of the noiseMemes or jokes: 7,308 views, 0.1% of the noise's views1.2%15 posts0.1% of viewsSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
About a fifth of posts are meta, and memes and hot takes are under half a percent each1,265 noise posts by sub-type. Audited: meta 21% (19 to 22) of all posts. Unrelated or unclear: 523 posts (41.3%), 33.1% of the noise's views. Benchmarks of the model: 360 posts (28.5%), 33.0% of the noise's views. Tooling or wrappers: 245 posts (19.4%), 19.6% of the noise's views. Explainers or tutorials: 59 posts (4.7%), 1.7% of the noise's views. News or reposts: 45 posts (3.6%), 10.6% of the noise's views. Hot takes: 18 posts (1.4%), 1.9% of the noise's views. Memes or jokes: 15 posts (1.2%), 0.1% of the noise's views.About a fifth of posts are meta, andmemes and hot takes are under half apercent each1,265 noise posts by sub-type. Audited: meta 21%(19 to 22) of all posts.0%25%50%Share of the noiseShare of the noise postsShare of the noise viewsUnrelated or unclearUnrelated or unclear: 523 posts, 41.3% of the noiseUnrelated or unclear: 2,176,307 views, 33.1% of the noise's views41.3%523 posts33.1% of viewsBenchmarks of the modelBenchmarks of the model: 360 posts, 28.5% of the noiseBenchmarks of the model: 2,169,703 views, 33.0% of the noise's views28.5%360 posts33.0% of viewsTooling or wrappersTooling or wrappers: 245 posts, 19.4% of the noiseTooling or wrappers: 1,291,340 views, 19.6% of the noise's views19.4%245 posts19.6% of viewsExplainers or tutorialsExplainers or tutorials: 59 posts, 4.7% of the noiseExplainers or tutorials: 113,146 views, 1.7% of the noise's views4.7%59 posts1.7% of viewsNews or repostsNews or reposts: 45 posts, 3.6% of the noiseNews or reposts: 695,749 views, 10.6% of the noise's views3.6%45 posts10.6% of viewsHot takesHot takes: 18 posts, 1.4% of the noiseHot takes: 123,574 views, 1.9% of the noise's views1.4%18 posts1.9% of viewsMemes or jokesMemes or jokes: 15 posts, 1.2% of the noiseMemes or jokes: 7,308 views, 0.1% of the noise's views1.2%15 posts0.1% of viewsSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
Sub-typePostsShare of noiseShare of all postsShare of noise viewsMedian viewsMost-viewed post (views)
Unrelated or unclear52341.3%9.3%33.1%137LLM built from 29 yes/no questions per character with Jev (158,692)
Benchmarks of the model36028.5%6.4%33.0%66Word-choice interface for Jev to generate text (649,221)
Tooling or wrappers24519.4%4.4%19.6%116Probabilistic programming language powered by Jev (774,848)
Explainers or tutorials594.7%1.1%1.7%114Dice game demo about calibrated probabilities (28,802)
News or reposts453.6%0.8%10.6%123Typed decision API for app state on Venice (527,680)
Hot takes181.4%0.3%1.9%137SEO/geo audit and fix agent cut costs 30x (58,987)
Memes or jokes151.2%0.3%0.1%123Tiny town website with Jev in the tallest tower (2,167)
All noise1,265100%22.6%100%111Probabilistic programming language powered by Jev (774,848)
Hot takes, bullish6
Hot takes, skeptical7
Hot takes, mixed4
Hot takes, neutral1

Figure 1c

What would have done the job before?

None of the 91 builds most likely to show something new, the posts labeled as a measured decision made while a person waits, did something that was unavailable before: a team would have used an LLM for 64, rules for 12, a vendor API for 11 and a classic model for 4. The audit judged each one from its post, not by running the build.

Of the 91 builds most likely to show something new, none did something that was unavailable before91 builds (85 distinct), by what a team would have used before Jev. An LLM: 64 of 91. Rules or heuristics: 12 of 91. A vendor API: 11 of 91. A classic model: 4 of 91. Unavailable at any price: 0 of 91.Of the 91 builds most likely to show something new, nonedid something that was unavailable before91 builds (85 distinct), by what a team would have used before Jev.An LLMAn LLM: 64 of 91 posts6470%Rules or heuristicsRules or heuristics: 12 of 91 posts1213%A vendor APIA vendor API: 11 of 91 posts1112%A classic modelA classic model: 4 of 91 posts44%Unavailable at any priceUnavailable at any price: 0 of 91 posts0of 91 postsSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
Of the 91 builds most likely to show something new, none did something that was unavailable before91 builds (85 distinct), by what a team would have used before Jev. An LLM: 64 of 91. Rules or heuristics: 12 of 91. A vendor API: 11 of 91. A classic model: 4 of 91. Unavailable at any price: 0 of 91.Of the 91 builds most likely to showsomething new, none did somethingthat was unavailable before91 builds (85 distinct), by what a team would haveused before Jev.An LLMAn LLM: 64 of 91 posts6470%Rules or heuristicsRules or heuristics: 12 of 91 posts1213%A vendor APIA vendor API: 11 of 91 posts1112%A classic modelA classic model: 4 of 91 posts44%Unavailable at any priceUnavailable at any price: 0 of 91 posts0of 91 postsSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
What would have done the job beforePostsShare of the candidates
An LLM6470%
Rules or heuristics1213%
A vendor API1112%
A classic model44%
Unavailable at any price00%
Distinct builds among them85
Job: feed or comment filters2224%
Job: typing or form fill1314%
Job: voice commands1314%
Job: triage67%
Job: voice turn-taking67%
Job: guards55%
Job: speech scoring55%
Job: trading decisions55%
Job: classification33%
Job: computer use33%
Job: games or simulations33%
Job: alerts22%
Job: live control22%
Job: search22%
Job: personal tools11%

Figure 2

Half the views went to 1 percent of posts

The top 1% of posts, 56 of them, hold 53.3% of the 44.6 million views and 47.0% of the likes, and the median post has 133 views. The most-viewed post alone is 7.8% of views, with a like rate far below the median, but without the 9 posts like it the top 1% still hold 49.8%.

Half the views went to 1 percent of posts: the top 56 held 53.3%5,595 posts, 44.6M views and 302,260 likes. Top 1% of posts: 53.3% of views, 47.0% of likes. Median post: 133 views.Half the views went to 1 percent of posts: the top 56 held53.3%5,595 posts, 44.6M views and 302,260 likes.ViewsLikesEqual share for every postCumulative share of views or likes0%25%50%75%100%0.01%0.1%1%10%100%Share of posts, most-viewed first (log scale)equal shareTop 0.1% of posts: 13.6% of likesTop 1% of posts: 47.0% of likesTop 5% of posts: 82.4% of likesTop 10% of posts: 91.7% of likesTop 25% of posts: 97.6% of likesTop 50% of posts: 99.5% of likesTop 0.1% of posts (6): 18.7% of viewsTop 1% of posts (56): 53.3% of viewsTop 5% of posts (280): 84.6% of viewsTop 10% of posts (560): 93.4% of viewsTop 25% of posts (1,399): 98.5% of viewsTop 50% of posts (2,798): 99.7% of viewsTop 1% of posts (56)53.3% of views47.0% of likesThe top post: 7.8%Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
Half the views went to 1 percent of posts: the top 56 held 53.3%5,595 posts, 44.6M views and 302,260 likes. Top 1% of posts: 53.3% of views, 47.0% of likes. Median post: 133 views.Half the views went to 1 percent ofposts: the top 56 held 53.3%5,595 posts, 44.6M views and 302,260 likes.ViewsLikesEqual share for every postCumulative share of views or likes0%25%50%75%100%0.01%0.1%1%10%100%Share of posts, most-viewed first (log scale)equal shareTop 0.1% of posts: 13.6% of likesTop 1% of posts: 47.0% of likesTop 5% of posts: 82.4% of likesTop 10% of posts: 91.7% of likesTop 25% of posts: 97.6% of likesTop 50% of posts: 99.5% of likesTop 0.1% of posts (6): 18.7% of viewsTop 1% of posts (56): 53.3% of viewsTop 5% of posts (280): 84.6% of viewsTop 10% of posts (560): 93.4% of viewsTop 25% of posts (1,399): 98.5% of viewsTop 50% of posts (2,798): 99.7% of viewsTop 1% of posts (56)53.3% of views47.0% of likesThe top post: 7.8%Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
Top share of postsPostsShare of viewsShare of likesSmallest post in group (views)
0.1%618.7%13.6%723,016
1%5653.3%47.0%149,377
5%28084.6%82.4%26,747
10%56093.4%91.7%7,202
25%1,39998.5%97.6%955
50%2,79899.7%99.5%133

Figure 3

What did they compare against?

About four in five posts compare Jev with nothing, 80% (77 to 82) on the audit, and about 1 in 30 with the classifiers, rules and vendor APIs it would replace, 3.7% (3 to 5). A week of "28× cheaper", and almost nobody asking "than what I already had?"

About four in five posts compared Jev with nothing, and about 1 in 30 with what it would replace5,595 posts. Audited: nothing 80% (77 to 82), the tools Jev would replace 3.7% (3 to 5). Nothing: 80.4%. Frontier LLM: 11.3%. Small LLM: 4.3%. Rules or regex: 2.5%. Classic classifier or ML: 1.1%. Vendor API: 0.6%. Audited: nothing 80% (77 to 82), frontier LLM 11.3% (10 to 13), small LLM 5.0% (4 to 7), the tools Jev would replace 3.7% (3 to 5).About four in five posts compared Jev with nothing, andabout 1 in 30 with what it would replace5,595 posts. Audited: nothing 80% (77 to 82), the tools Jev would replace 3.7% (3 to 5).0%25%50%75%100%Opus's labelsAudit estimate and 95% intervalNothingNothing: 80.4% of posts (4,498), Opus's labels80.4%4,498 postsNothing: audit 80.0% (95% interval 77.4% to 81.6%)audit 80% (77 to 82)Frontier LLMFrontier LLM: 11.3% of posts (630), Opus's labels11.3%630 postsFrontier LLM: audit 11.3% (95% interval 10.5% to 13.1%)audit 11.3% (10 to 13)Small LLMSmall LLM: 4.3% of posts (239), Opus's labels4.3%239 postsSmall LLM: audit 5.0% (95% interval 4.2% to 6.9%)audit 5.0% (4 to 7)Rules or regexRules or regex: 2.5% of posts (138), Opus's labels2.5%138 postsClassic classifier or MLClassic classifier or ML: 1.1% of posts (59), Opus's labels1.1%59 postsVendor APIVendor API: 0.6% of posts (31), Opus's labels0.6%31 posts4.1% combined (228 posts): the tools Jevwould replace. Audit: 3.7% (3 to 5)Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
About four in five posts compared Jev with nothing, and about 1 in 30 with what it would replace5,595 posts. Audited: nothing 80% (77 to 82), the tools Jev would replace 3.7% (3 to 5). Nothing: 80.4%. Frontier LLM: 11.3%. Small LLM: 4.3%. Rules or regex: 2.5%. Classic classifier or ML: 1.1%. Vendor API: 0.6%. Audited: nothing 80% (77 to 82), frontier LLM 11.3% (10 to 13), small LLM 5.0% (4 to 7), the tools Jev would replace 3.7% (3 to 5).About four in five posts comparedJev with nothing, and about 1 in 30with what it would replace5,595 posts. Audited: nothing 80% (77 to 82), thetools Jev would replace 3.7% (3 to 5).0%25%50%75%100%Opus's labelsAudit estimate and 95% intervalNothingNothing: 80.4% of posts (4,498), Opus's labels80.4%4,498 postsNothing: audit 80.0% (95% interval 77.4% to 81.6%)audit 80% (77 to 82)Frontier LLMFrontier LLM: 11.3% of posts (630), Opus's labels11.3%630 postsFrontier LLM: audit 11.3% (95% interval 10.5% to 13.1%)audit 11.3% (10 to 13)Small LLMSmall LLM: 4.3% of posts (239), Opus's labels4.3%239 postsSmall LLM: audit 5.0% (95% interval 4.2% to 6.9%)audit 5.0% (4 to 7)Rules or regexRules or regex: 2.5% of posts (138), Opus's labels2.5%138 postsClassic classifier or MLClassic classifier or ML: 1.1% of posts (59), Opus's labels1.1%59 postsVendor APIVendor API: 0.6% of posts (31), Opus's labels0.6%31 posts4.1% combined (228posts): the tools Jevwould replace. Audit:3.7% (3 to 5)Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
BaselinePostsShare of postsShare of viewsAudit estimate (95% interval)
Nothing4,49880.4%74.9%80% (77 to 82)
Frontier LLM63011.3%14.3%11.3% (10 to 13)
Small LLM2394.3%6.9%5.0% (4 to 7)
Rules or regex1382.5%2.2%
Classic classifier or ML591.1%0.5%
Vendor API310.6%1.3%
The tools Jev would replace (the last three)2284.1%3.7% (3 to 5)

Figure 4

Claims without numbers

By family, data and telemetry measured most often, 65% of its posts, and collaboration and typing least, 17%. Only eight posts measured Jev in production on the audit's re-read: a task router deployed "for a real production use case", a deploy approval gate, a news pipeline, a search reranker, an episode recommender, a fleet-wide decision layer, and two public Q&A sites, AskJev.ai and AskJev.net, that counted their questions and visitors.

Two thirds of builds showed no numbers, and 8 of 5,595 measured production on a blind re-read5,595 posts. Audited: no measurement 67% (64 to 70), production 0.14% (at most 2%). Labeled production: 37 posts. Measured demo: 1,723 posts. Demo, no numbers: 3,672 posts. Proposal or idea: 21 posts. Commentary or meme: 142 posts.Two thirds of builds showed no numbers, and 8 of 5,595measured production on a blind re-read5,595 posts. Audited: no measurement 67% (64 to 70), production 0.14% (at most 2%).0%50%100%Share of the family’s postsLabeled productionMeasured demoDemo, no numbersProposal or ideaCommentary or mememeasuredAll posts5,595All posts, labeled production: 37 posts (0.7%)All posts, measured demo: 1,723 posts (30.8%)All posts, demo, no numbers: 3,672 posts (65.6%)All posts, proposal or idea: 21 posts (0.4%)All posts, commentary or meme: 142 posts (2.5%)31%Data and telemetry271Data and telemetry, labeled production: 2 posts (0.7%)Data and telemetry, measured demo: 173 posts (63.8%)Data and telemetry, demo, no numbers: 94 posts (34.7%)Data and telemetry, commentary or meme: 2 posts (0.7%)65%Trading and markets268Trading and markets, labeled production: 1 posts (0.4%)Trading and markets, measured demo: 122 posts (45.5%)Trading and markets, demo, no numbers: 143 posts (53.4%)Trading and markets, proposal or idea: 1 posts (0.4%)Trading and markets, commentary or meme: 1 posts (0.4%)46%Compaction and context87Compaction and context, measured demo: 35 posts (40.2%)Compaction and context, demo, no numbers: 50 posts (57.5%)Compaction and context, proposal or idea: 1 posts (1.1%)Compaction and context, commentary or meme: 1 posts (1.1%)40%Classification and routing781Classification and routing, labeled production: 10 posts (1.3%)Classification and routing, measured demo: 304 posts (38.9%)Classification and routing, demo, no numbers: 457 posts (58.5%)Classification and routing, proposal or idea: 6 posts (0.8%)Classification and routing, commentary or meme: 4 posts (0.5%)40%Search, rerank, extraction518Search, rerank, extraction, labeled production: 8 posts (1.5%)Search, rerank, extraction, measured demo: 195 posts (37.6%)Search, rerank, extraction, demo, no numbers: 313 posts (60.4%)Search, rerank, extraction, proposal or idea: 2 posts (0.4%)39%Browser and computer use254Browser and computer use, measured demo: 95 posts (37.4%)Browser and computer use, demo, no numbers: 159 posts (62.6%)37%Moderation and guardrails257Moderation and guardrails, labeled production: 3 posts (1.2%)Moderation and guardrails, measured demo: 78 posts (30.4%)Moderation and guardrails, demo, no numbers: 171 posts (66.5%)Moderation and guardrails, proposal or idea: 4 posts (1.6%)Moderation and guardrails, commentary or meme: 1 posts (0.4%)32%Agent harness and gating262Agent harness and gating, labeled production: 2 posts (0.8%)Agent harness and gating, measured demo: 77 posts (29.4%)Agent harness and gating, demo, no numbers: 182 posts (69.5%)Agent harness and gating, proposal or idea: 1 posts (0.4%)30%Evals and judging366Evals and judging, labeled production: 3 posts (0.8%)Evals and judging, measured demo: 106 posts (29.0%)Evals and judging, demo, no numbers: 255 posts (69.7%)Evals and judging, proposal or idea: 1 posts (0.3%)Evals and judging, commentary or meme: 1 posts (0.3%)30%Voice and turn-taking72Voice and turn-taking, measured demo: 21 posts (29.2%)Voice and turn-taking, demo, no numbers: 51 posts (70.8%)29%Games and control loops1,071Games and control loops, measured demo: 280 posts (26.1%)Games and control loops, demo, no numbers: 789 posts (73.7%)Games and control loops, proposal or idea: 1 posts (0.1%)Games and control loops, commentary or meme: 1 posts (0.1%)26%Other or meta1,254Other or meta, labeled production: 7 posts (0.6%)Other or meta, measured demo: 215 posts (17.1%)Other or meta, demo, no numbers: 897 posts (71.5%)Other or meta, proposal or idea: 4 posts (0.3%)Other or meta, commentary or meme: 131 posts (10.4%)18%Live chat and streams68Live chat and streams, labeled production: 1 posts (1.5%)Live chat and streams, measured demo: 11 posts (16.2%)Live chat and streams, demo, no numbers: 56 posts (82.4%)18%Collaboration and typing66Collaboration and typing, measured demo: 11 posts (16.7%)Collaboration and typing, demo, no numbers: 55 posts (83.3%)17%Measured production: 8 posts on the audit's re-readSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
Two thirds of builds showed no numbers, and 8 of 5,595 measured production on a blind re-read5,595 posts. Audited: no measurement 67% (64 to 70), production 0.14% (at most 2%). Labeled production: 37 posts. Measured demo: 1,723 posts. Demo, no numbers: 3,672 posts. Proposal or idea: 21 posts. Commentary or meme: 142 posts.Two thirds of builds showed nonumbers, and 8 of 5,595 measuredproduction on a blind re-read5,595 posts. Audited: no measurement 67% (64 to70), production 0.14% (at most 2%).0%50%100%Share of the family’s postsLabeled productionMeasured demoDemo, no numbersProposal or ideaCommentary or mememeasuredAll posts5,595All posts, labeled production: 37 posts (0.7%)All posts, measured demo: 1,723 posts (30.8%)All posts, demo, no numbers: 3,672 posts (65.6%)All posts, proposal or idea: 21 posts (0.4%)All posts, commentary or meme: 142 posts (2.5%)31%Data and telemetry271Data and telemetry, labeled production: 2 posts (0.7%)Data and telemetry, measured demo: 173 posts (63.8%)Data and telemetry, demo, no numbers: 94 posts (34.7%)Data and telemetry, commentary or meme: 2 posts (0.7%)65%Trading and markets268Trading and markets, labeled production: 1 posts (0.4%)Trading and markets, measured demo: 122 posts (45.5%)Trading and markets, demo, no numbers: 143 posts (53.4%)Trading and markets, proposal or idea: 1 posts (0.4%)Trading and markets, commentary or meme: 1 posts (0.4%)46%Compaction and context87Compaction and context, measured demo: 35 posts (40.2%)Compaction and context, demo, no numbers: 50 posts (57.5%)Compaction and context, proposal or idea: 1 posts (1.1%)Compaction and context, commentary or meme: 1 posts (1.1%)40%Classification and routing781Classification and routing, labeled production: 10 posts (1.3%)Classification and routing, measured demo: 304 posts (38.9%)Classification and routing, demo, no numbers: 457 posts (58.5%)Classification and routing, proposal or idea: 6 posts (0.8%)Classification and routing, commentary or meme: 4 posts (0.5%)40%Search, rerank, extraction518Search, rerank, extraction, labeled production: 8 posts (1.5%)Search, rerank, extraction, measured demo: 195 posts (37.6%)Search, rerank, extraction, demo, no numbers: 313 posts (60.4%)Search, rerank, extraction, proposal or idea: 2 posts (0.4%)39%Browser and computer use254Browser and computer use, measured demo: 95 posts (37.4%)Browser and computer use, demo, no numbers: 159 posts (62.6%)37%Moderation and guardrails257Moderation and guardrails, labeled production: 3 posts (1.2%)Moderation and guardrails, measured demo: 78 posts (30.4%)Moderation and guardrails, demo, no numbers: 171 posts (66.5%)Moderation and guardrails, proposal or idea: 4 posts (1.6%)Moderation and guardrails, commentary or meme: 1 posts (0.4%)32%Agent harness and gating262Agent harness and gating, labeled production: 2 posts (0.8%)Agent harness and gating, measured demo: 77 posts (29.4%)Agent harness and gating, demo, no numbers: 182 posts (69.5%)Agent harness and gating, proposal or idea: 1 posts (0.4%)30%Evals and judging366Evals and judging, labeled production: 3 posts (0.8%)Evals and judging, measured demo: 106 posts (29.0%)Evals and judging, demo, no numbers: 255 posts (69.7%)Evals and judging, proposal or idea: 1 posts (0.3%)Evals and judging, commentary or meme: 1 posts (0.3%)30%Voice and turn-taking72Voice and turn-taking, measured demo: 21 posts (29.2%)Voice and turn-taking, demo, no numbers: 51 posts (70.8%)29%Games and control loops1,071Games and control loops, measured demo: 280 posts (26.1%)Games and control loops, demo, no numbers: 789 posts (73.7%)Games and control loops, proposal or idea: 1 posts (0.1%)Games and control loops, commentary or meme: 1 posts (0.1%)26%Other or meta1,254Other or meta, labeled production: 7 posts (0.6%)Other or meta, measured demo: 215 posts (17.1%)Other or meta, demo, no numbers: 897 posts (71.5%)Other or meta, proposal or idea: 4 posts (0.3%)Other or meta, commentary or meme: 131 posts (10.4%)18%Live chat and streams68Live chat and streams, labeled production: 1 posts (1.5%)Live chat and streams, measured demo: 11 posts (16.2%)Live chat and streams, demo, no numbers: 56 posts (82.4%)18%Collaboration and typing66Collaboration and typing, measured demo: 11 posts (16.7%)Collaboration and typing, demo, no numbers: 55 posts (83.3%)17%Measured production: 8 posts on the audit's re-readSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
FamilyLabeled productionMeasured demoDemo, no numbersProposal or ideaCommentary or memeMeasured share
All posts371,7233,6722114231.5%
Data and telemetry2173940264.6%
Trading and markets11221431145.9%
Compaction and context035501140.2%
Classification and routing103044576440.2%
Search, rerank, extraction81953132039.2%
Browser and computer use0951590037.4%
Moderation and guardrails3781714131.5%
Agent harness and gating2771821030.2%
Evals and judging31062551129.8%
Voice and turn-taking021510029.2%
Games and control loops02807891126.1%
Other or meta7215897413117.7%
Live chat and streams111560017.6%
Collaboration and typing011550016.7%

Figure 5

What people claimed

The median claim on the cards was 28× cheaper and 6× faster, close to OpenChamber's own survey of user reports at about 30× and 7×. These are the authors' numbers, and figure 3 shows what most of them are measured against: nothing named at all, or a frontier model.

The typical claim was 28× cheaper and 6× faster, in line with OpenChamber's own survey126 cost chips on 120 posts and 297 speed chips on 281 posts. Cost claims, N× cheaper: median 28×, interquartile 5.25× to 100×. Speed claims, N× faster: median 6×, interquartile 2× to 16×.The typical claim was 28× cheaper and 6× faster, in linewith OpenChamber's own survey126 cost chips on 120 posts and 297 speed chips on 281 posts.Cost claims, N× cheaper126 chipsmedian 28× (34× without 1× chips)OpenChamber survey, about 30×1× chipsvs small LLMs on the same question (Pong, measured from Vercel)0102030Exactly 1×: 8 chips1× to 2×: 1 chips2× to 5×: 19 chips5× to 10×: 14 chips10× to 20×: 12 chips20× to 50×: 24 chips50× to 100×: 10 chips100× to 200×: 16 chips200× to 500×: 15 chips500× to 1,000×: 3 chips1,000× to more: 4 chips<11251020501002005001kSpeed claims, N× faster297 chipsmedian 6× (8.1× without 1× chips)OpenChamber survey, about 7×1× chipsvs small LLMs on the same question (Pong, measured from Vercel)0255075under 1×: 3 chipsExactly 1×: 51 chips1× to 2×: 11 chips2× to 5×: 68 chips5× to 10×: 51 chips10× to 20×: 47 chips20× to 50×: 29 chips50× to 100×: 8 chips100× to 200×: 16 chips200× to 500×: 6 chips500× to 1,000×: 5 chips1,000× to more: 2 chips<11251020501002005001kClaimed multiple (×), log bins; bars count chipsSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
The typical claim was 28× cheaper and 6× faster, in line with OpenChamber's own survey126 cost chips on 120 posts and 297 speed chips on 281 posts. Cost claims, N× cheaper: median 28×, interquartile 5.25× to 100×. Speed claims, N× faster: median 6×, interquartile 2× to 16×.The typical claim was 28× cheaperand 6× faster, in line withOpenChamber's own survey126 cost chips on 120 posts and 297 speed chips on281 posts.Cost claims, N× cheaper126 chipsmedian 28× (34× without 1× chips)OpenChamber survey, about 30×1× chipsvs small LLMs on the same question (Pong, measured from Vercel)0102030Exactly 1×: 8 chips1× to 2×: 1 chips2× to 5×: 19 chips5× to 10×: 14 chips10× to 20×: 12 chips20× to 50×: 24 chips50× to 100×: 10 chips100× to 200×: 16 chips200× to 500×: 15 chips500× to 1,000×: 3 chips1,000× to more: 4 chips<11251020501002005001kSpeed claims, N× faster297 chipsmedian 6× (8.1× without 1× chips)OpenChamber survey, about 7×1× chipsvs small LLMs on the same question (Pong, measured from Vercel)0255075under 1×: 3 chipsExactly 1×: 51 chips1× to 2×: 11 chips2× to 5×: 68 chips5× to 10×: 51 chips10× to 20×: 47 chips20× to 50×: 29 chips50× to 100×: 8 chips100× to 200×: 16 chips200× to 500×: 6 chips500× to 1,000×: 5 chips1,000× to more: 2 chips<11251020501002005001kClaimed multiple (×), log bins; bars count chipsSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
Claimed multipleCost claims, N× cheaper (chips)Speed claims, N× faster (chips)
under 1×03
1× to under 2×9 (8 read exactly 1×)62 (51 read exactly 1×)
2× to under 5×1968
5× to under 10×1451
10× to under 20×1247
20× to under 50×2429
50× to under 100×108
100× to under 200×1616
200× to under 500×156
500× to under 1,000×35
1,000× to more42
Median28× (34× without 1×)6× (8.1× without 1×)
Middle half5.25× to 100×2× to 16×

Figure 6

Who needs it fast

On the audit, about 11% of posts need a decision in under 300 ms, and about nine in ten of those are games. Jev's own median call took 412 ms from a laptop, above every one of those budgets.

About 11 percent of posts need a decision in under 300 ms, and about 9 in 10 of those are games5,595 posts. Audited: under 300 ms 11.0% (9.5 to 12.5), games 91% of those (82 to 95). Under 300 ms: 11.0% of posts (9.5 to 12.5); games 90.8% of those.About 11 percent of posts need a decision in under 300ms, and about 9 in 10 of those are games5,595 posts. Audited: under 300 ms 11.0% (9.5 to 12.5), games 91% of those (82 to 95).Under 300 ms: 11.0% of posts, 95% interval 9.5% to 12.5%Games under 300 ms: 10.0% of postsOther posts under 300 ms: 1.0% of postsEverything else: 89.0% of posts0%25%50%75%100%Share of the 5,595 postsGames under 300 msOther posts under 300 msEverything else95% intervalSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
About 11 percent of posts need a decision in under 300 ms, and about 9 in 10 of those are games5,595 posts. Audited: under 300 ms 11.0% (9.5 to 12.5), games 91% of those (82 to 95). Under 300 ms: 11.0% of posts (9.5 to 12.5); games 90.8% of those.About 11 percent of posts need adecision in under 300 ms, and about 9in 10 of those are games5,595 posts. Audited: under 300 ms 11.0% (9.5 to12.5), games 91% of those (82 to 95).Under 300 ms: 11.0% of posts, 95% interval 9.5% to 12.5%Games under 300 ms: 10.0% of postsOther posts under 300 ms: 1.0% of postsEverything else: 89.0% of posts0%25%50%75%100%Share of the 5,595 postsGames under 300 msOther posts under 300 msEverything else95% intervalSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
Audited estimateEstimate95% intervalFrom
Posts that need a decision in under 300 ms11.0%9.5% to 12.5%76 of 150 re-read posts still under 300 ms, plus 91 that only Opus puts there
The same, on the audit's samples alone10.3%8.2% to 14.3%neither model's word
Games, among those posts90.8%82.2% to 95.5%69 of 76

Figure 7

The realtime slice

Outside games, a live loop is 8.1% of posts on the audit (6.0 to 11.4), and none of the eight posts that measured production is a realtime build. The best realtime demos, a voice turn-end detector at 0.224 s and a predictive spreadsheet that scores rows as you type, are faster versions of jobs a vendor API and an LLM already did.

Outside games, about 8 percent of posts have a live loop, and none of them is in production376 posts flagged realtime outside games. Audited: a live loop 8.1% of posts (6.0 to 11.4). Labeled production: 2. Measured demo: 123. Demo, no numbers: 250. Proposal or commentary: 1.Outside games, about 8 percent of posts have a live loop,and none of them is in production376 posts flagged realtime outside games. Audited: a live loop 8.1% of posts (6.0 to 11.4).020406080100120Posts Opus flags realtimeLabeled production 2Measured demo 123Demo, no numbers 250Proposal or commentary 1Trading and marketsTrading and markets, labeled production: 1 postsTrading and markets, measured demo: 60 postsTrading and markets, demo, no numbers: 43 posts104Voice and turn-takingVoice and turn-taking, measured demo: 21 postsVoice and turn-taking, demo, no numbers: 48 posts69Live chat and streamsLive chat and streams, labeled production: 1 postsLive chat and streams, measured demo: 9 postsLive chat and streams, demo, no numbers: 54 posts64Collaboration and typingCollaboration and typing, measured demo: 11 postsCollaboration and typing, demo, no numbers: 51 posts62Browser and computer useBrowser and computer use, measured demo: 3 postsBrowser and computer use, demo, no numbers: 16 posts19Moderation and guardrailsModeration and guardrails, measured demo: 9 postsModeration and guardrails, demo, no numbers: 5 postsModeration and guardrails, proposal or commentary: 1 posts15Other or metaOther or meta, measured demo: 1 postsOther or meta, demo, no numbers: 14 posts15Classification and routingClassification and routing, measured demo: 4 postsClassification and routing, demo, no numbers: 5 posts9Data and telemetryData and telemetry, demo, no numbers: 7 posts7Search, rerank, extractionSearch, rerank, extraction, measured demo: 3 postsSearch, rerank, extraction, demo, no numbers: 3 posts6Evals and judgingEvals and judging, measured demo: 2 postsEvals and judging, demo, no numbers: 2 posts4Agent harness and gatingAgent harness and gating, demo, no numbers: 2 posts2Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
Outside games, about 8 percent of posts have a live loop, and none of them is in production376 posts flagged realtime outside games. Audited: a live loop 8.1% of posts (6.0 to 11.4). Labeled production: 2. Measured demo: 123. Demo, no numbers: 250. Proposal or commentary: 1.Outside games, about 8 percent ofposts have a live loop, and none ofthem is in production376 posts flagged realtime outside games. Audited:a live loop 8.1% of posts (6.0 to 11.4).020406080100120Posts Opus flags realtimeLabeled production 2Measured demo 123Demo, no numbers 250Proposal or commentary 1Trading and marketsTrading and markets, labeled production: 1 postsTrading and markets, measured demo: 60 postsTrading and markets, demo, no numbers: 43 posts104Voice and turn-takingVoice and turn-taking, measured demo: 21 postsVoice and turn-taking, demo, no numbers: 48 posts69Live chat and streamsLive chat and streams, labeled production: 1 postsLive chat and streams, measured demo: 9 postsLive chat and streams, demo, no numbers: 54 posts64Collaboration and typingCollaboration and typing, measured demo: 11 postsCollaboration and typing, demo, no numbers: 51 posts62Browser and computer useBrowser and computer use, measured demo: 3 postsBrowser and computer use, demo, no numbers: 16 posts19Moderation and guardrailsModeration and guardrails, measured demo: 9 postsModeration and guardrails, demo, no numbers: 5 postsModeration and guardrails, proposal or commentary: 1 posts15Other or metaOther or meta, measured demo: 1 postsOther or meta, demo, no numbers: 14 posts15Classification and routingClassification and routing, measured demo: 4 postsClassification and routing, demo, no numbers: 5 posts9Data and telemetryData and telemetry, demo, no numbers: 7 posts7Search, rerank, extractionSearch, rerank, extraction, measured demo: 3 postsSearch, rerank, extraction, demo, no numbers: 3 posts6Evals and judgingEvals and judging, measured demo: 2 postsEvals and judging, demo, no numbers: 2 posts4Agent harness and gatingAgent harness and gating, demo, no numbers: 2 posts2Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
FamilyLabeled productionMeasured demoDemo, no numbersProposal or commentaryFlagged realtime
Trading and markets160430104
Voice and turn-taking02148069
Live chat and streams1954064
Collaboration and typing01151062
Browser and computer use0316019
Moderation and guardrails095115
Other or meta0114015
Classification and routing04509
Data and telemetry00707
Search, rerank, extraction03306
Evals and judging02204
Agent harness and gating00202
All outside games21232501376

Figure 8

What measuring buys

Measuring buys the hits, not the typical post: measured demos averaged 12,755 views against 5,752 for demos with no numbers, but the median barely moves, 143 against 130. Most posts got little attention either way: 45% had under 100 views.

Measured demos averaged twice the views of demos with no numbers, but the typical post gained little5,558 posts by evidence level, without the 37 labeled production. Measured demo: 1,723 posts, median 143, mean 12,755 views. Demo, no numbers: 3,672 posts, median 130, mean 5,752 views. Proposal or idea: 21 posts, median 55, mean 305 views. Commentary or meme: 142 posts, median 119, mean 7,246 views.Measured demos averaged twice the views of demoswith no numbers, but the typical post gained little5,558 posts by evidence level, without the 37 labeled production.050100150Median views per postMeasured demo1,723 postsMeasured demo: median 143 views per post, 1,723 posts143Demo, no numbers3,672 postsDemo, no numbers: median 130 views per post, 3,672 posts130Proposal or idea21 postsProposal or idea: median 55 views per post, 21 posts55Commentary or meme142 postsCommentary or meme: median 119 views per post, 142 posts11905k10k15kMean views per postMeasured demo1,723 postsMeasured demo: mean 12,755 views per post, 1,723 posts12,755Demo, no numbers3,672 postsDemo, no numbers: mean 5,752 views per post, 3,672 posts5,752Proposal or idea21 postsProposal or idea: mean 305 views per post, 21 posts305Commentary or meme142 postsCommentary or meme: mean 7,246 views per post, 142 posts7,246Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
Measured demos averaged twice the views of demos with no numbers, but the typical post gained little5,558 posts by evidence level, without the 37 labeled production. Measured demo: 1,723 posts, median 143, mean 12,755 views. Demo, no numbers: 3,672 posts, median 130, mean 5,752 views. Proposal or idea: 21 posts, median 55, mean 305 views. Commentary or meme: 142 posts, median 119, mean 7,246 views.Measured demos averaged twice theviews of demos with no numbers, butthe typical post gained little5,558 posts by evidence level, without the 37labeled production.050100150Median views per postMeasured demo1,723 postsMeasured demo: median 143 views per post, 1,723 posts143Demo, no numbers3,672 postsDemo, no numbers: median 130 views per post, 3,672 posts130Proposal or idea21 postsProposal or idea: median 55 views per post, 21 posts55Commentary or meme142 postsCommentary or meme: median 119 views per post, 142 posts11905k10k15kMean views per postMeasured demo1,723 postsMeasured demo: mean 12,755 views per post, 1,723 posts12,755Demo, no numbers3,672 postsDemo, no numbers: mean 5,752 views per post, 3,672 posts5,752Proposal or idea21 postsProposal or idea: mean 305 views per post, 21 posts305Commentary or meme142 postsCommentary or meme: mean 7,246 views per post, 142 posts7,246Source: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
EvidencePostsMedian viewsMean viewsShare of views
Measured demo1,72314312,75549.3%
Demo, no numbers3,6721305,75247.4%
Proposal or idea21553050.0%
Commentary or meme1421197,2462.3%

Figure 9

Jev grading Jev

When Jev was surest it picked Opus's family 98.0% of the time, against 42.5% when it was least sure, and it labeled every post for $0.33 against $29.26 for Opus: a cheap model that knows when it's sure can sit in front of a slower one. That's agreement, not accuracy, but on the 464 audited posts where Jev said 0.99 or more, it matched the audit's family 98.1% of the time.

When Jev was surest, it picked the same family as Opus 98.0% of the time5,949 Jev calls in tenths. Audited: at 0.99 or more, Jev matched 98.1% of 464 posts. D1: 42.5%. D2: 49.1%. D3: 65.2%. D4: 76.5%. D5: 83.2%. D6: 86.9%. D7: 90.1%. D8: 95.8%. D9: 97.1%. D10: 98.0%.When Jev was surest, it picked the same family as Opus98.0% of the time5,949 Jev calls in tenths. Audited: at 0.99 or more, Jev matched 98.1% of 464 posts.Jev agreed with OpusDisagreedJev’s average stated probabilityCards0200400600D1 (stated 0.20–0.52): agreed on 253 of 595, 42.5%D1: disagreed on 342 of 59542.5%D10.20–0.52D2 (stated 0.52–0.64): agreed on 292 of 595, 49.1%D2: disagreed on 303 of 595D20.52–0.64D3 (stated 0.64–0.75): agreed on 388 of 595, 65.2%D3: disagreed on 207 of 595D30.64–0.75D4 (stated 0.75–0.84): agreed on 455 of 595, 76.5%D4: disagreed on 140 of 595D40.75–0.84D5 (stated 0.84–0.91): agreed on 495 of 595, 83.2%D5: disagreed on 100 of 59583.2%D50.84–0.91D6 (stated 0.91–0.96): agreed on 516 of 594, 86.9%D6: disagreed on 78 of 594D60.91–0.96D7 (stated 0.96–0.98): agreed on 536 of 595, 90.1%D7: disagreed on 59 of 595D70.96–0.98D8 (stated 0.98–0.99): agreed on 570 of 595, 95.8%D8: disagreed on 25 of 595D80.98–0.99D9 (stated 0.99–1.00): agreed on 578 of 595, 97.1%D9: disagreed on 17 of 595D90.99–1.00D10 (stated 1.00–1.00): agreed on 583 of 595, 98.0%D10: disagreed on 12 of 59598.0%D101.00–1.00Tenths of Jev’s calls, least to most sure, with the stated probability rangeSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
When Jev was surest, it picked the same family as Opus 98.0% of the time5,949 Jev calls in tenths. Audited: at 0.99 or more, Jev matched 98.1% of 464 posts. D1: 42.5%. D2: 49.1%. D3: 65.2%. D4: 76.5%. D5: 83.2%. D6: 86.9%. D7: 90.1%. D8: 95.8%. D9: 97.1%. D10: 98.0%.When Jev was surest, it picked thesame family as Opus 98.0% of thetime5,949 Jev calls in tenths. Audited: at 0.99 or more,Jev matched 98.1% of 464 posts.Jev agreed with OpusDisagreedJev’s average stated probabilityCards0200400600D1 (stated 0.20–0.52): agreed on 253 of 595, 42.5%D1: disagreed on 342 of 59542.5%D1D2 (stated 0.52–0.64): agreed on 292 of 595, 49.1%D2: disagreed on 303 of 595D2D3 (stated 0.64–0.75): agreed on 388 of 595, 65.2%D3: disagreed on 207 of 595D3D4 (stated 0.75–0.84): agreed on 455 of 595, 76.5%D4: disagreed on 140 of 595D4D5 (stated 0.84–0.91): agreed on 495 of 595, 83.2%D5: disagreed on 100 of 59583.2%D5D6 (stated 0.91–0.96): agreed on 516 of 594, 86.9%D6: disagreed on 78 of 594D6D7 (stated 0.96–0.98): agreed on 536 of 595, 90.1%D7: disagreed on 59 of 595D7D8 (stated 0.98–0.99): agreed on 570 of 595, 95.8%D8: disagreed on 25 of 595D8D9 (stated 0.99–1.00): agreed on 578 of 595, 97.1%D9: disagreed on 17 of 595D9D10 (stated 1.00–1.00): agreed on 583 of 595, 98.0%D10: disagreed on 12 of 59598.0%D10Tenths of Jev’s calls, least to most sureSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
TenthCardsStated probabilityMean statedAgreement with Claude Opus 5.5
D15950.20 to 0.520.44042.5%
D25950.52 to 0.640.57949.1%
D35950.64 to 0.750.69465.2%
D45950.75 to 0.840.79776.5%
D55950.84 to 0.910.87783.2%
D65940.91 to 0.960.93386.9%
D75950.96 to 0.980.97090.1%
D85950.98 to 0.990.98895.8%
D95950.99 to 1.000.99997.1%
D105951.00 to 1.001.00098.0%

Figure 10

Where the builders were

Japanese builders wrote 20% of all posts but 32% to 36% of the live chat, voice and collaboration posts, the families I care about most. The feed starts at 00:28 UTC on the 16th, about six hours after launch, so the first evening is missing.

One post in five was in Japanese, and posting peaked on day 45,595 posts by language and by posting day (UTC). English: 68.7%. Japanese: 19.6%. Chinese: 6.5%. 25 other language codes: 5.2%. Peak day 2026-09-19: 1,027 posts.One post in five was in Japanese, and posting peaked onday 45,595 posts by language and by posting day (UTC).Language of the postEnglishEnglish: 3,846 posts (68.7%)68.7%3,846JapaneseJapanese: 1,098 posts (19.6%)19.6%1,098ChineseChinese: 361 posts (6.5%)6.5%36125 other language codes25 other language codes: 290 posts (5.2%)5.2%290Posts per day (UTC), September02505007501,0001,2502026-09-16: 89 posts (partial day)8916partial2026-09-17: 721 posts172026-09-18: 948 posts182026-09-19: 1,027 posts1,027192026-09-20: 945 posts202026-09-21: 901 posts212026-09-22: 628 posts222026-09-23: 336 posts (partial day)33623partialSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep 2026. Method and audit: see the end of this page.
One post in five was in Japanese, and posting peaked on day 45,595 posts by language and by posting day (UTC). English: 68.7%. Japanese: 19.6%. Chinese: 6.5%. 25 other language codes: 5.2%. Peak day 2026-09-19: 1,027 posts.One post in five was in Japanese,and posting peaked on day 45,595 posts by language and by posting day (UTC).Language of the postEnglishEnglish: 3,846 posts (68.7%)68.7%3,846JapaneseJapanese: 1,098 posts (19.6%)19.6%1,098ChineseChinese: 361 posts (6.5%)6.5%36125 other language codes25 other language codes: 290 posts (5.2%)5.2%290Posts per day (UTC), September02505007501,0001,2502026-09-16: 89 posts (partial day)8916partial2026-09-17: 721 posts172026-09-18: 948 posts182026-09-19: 1,027 posts1,027192026-09-20: 945 posts202026-09-21: 901 posts212026-09-22: 628 posts222026-09-23: 336 posts (partial day)33623partialSource: jev.openchamber.dev, 5,950 posts, 16 to 23 Sep2026. Method and audit: see the end of this page.
Show the numbers
Language or dayPostsShare of posts
English3,84668.7%
Japanese1,09819.6%
Chinese3616.5%
25 other language codes2905.2%
2026-09-16 (partial)891.6%
2026-09-1772112.9%
2026-09-1894816.9%
2026-09-191,02718.4%
2026-09-2094516.9%
2026-09-2190116.1%
2026-09-2262811.2%
2026-09-23 (partial)3366.0%

Method

How this was measured

The posts come from OpenChamber's Jev feed, snapshot 23 September 18:47 UTC: 5,950 posts from 4,593 authors between 16 and 23 September, as OpenChamber selected them. Without 7 duplicates, the 347 posts that don't use Jev and the one post Opus refused to label, 5,595 remain.

A written rubric sorted each post into one of fourteen families by what Jev decides, and labeled what it measured, what it compared Jev with and whether it needs realtime infrastructure.

Claude Sonnet 5 labeled every post first. A critical review of those labels tightened the rubric, and Sonnet labeled every post again. Then Grok, an independent model, labeled 1,681 posts blind: every post in the rare groups the headlines rest on, and random samples of the rest. Sonnet agreed with it on what a post is for only 71% of the time, so Claude Opus 5.5 labeled every post a final time; it agrees 82% of the time, and the charts use its labels.

That's why the headlines come with ranges: the audit read samples, so each share it checked is an estimate with a 95% interval, and the page uses it, not the labels' count.

Two fields are dropped. Framing, what a post leads with, misrepresented the many posts that lead with cost and latency together, and the seven-level latency tier agrees with the audit least (75%), so figure 6 keeps only the split at 300 ms.

I labeled 13 posts myself as a calibration, too few to check the models but enough to find both problems: I disagreed with both models on the tier of 8.

At Vercel AI Gateway list prices the labeling cost $47.53: $16.47 for the two Sonnet passes, $29.26 for Opus, $1.47 to sort the noise and $0.33 for Jev.

The labels, tables, rubric and code are at github.com/mattheworiordan/jev-landscape, without the text of any post: code under MIT, data and method under CC BY 4.0. This page and its charts are © 2026 Matthew O'Riordan; the posts belong to their authors.

Disclosure. I'm CEO of Ably, a realtime infrastructure company.

Show the audit tables

The audit

A piece about unmeasured claims should show its own measurements being checked. An independent reviewer, Grok, labeled 1,681 posts against the same rubric without seeing the model's labels: all 422 posts in the rare groups the headlines rest on, and seeded random samples of 100 to 300 posts from the rest. Each sample is scaled up to its group with a 95% Wilson interval. The review is published with every label and the scripts that turn them into the numbers on this page. The audit sampled from the Sonnet 5 labels. Because the charts use Opus's, I reran its estimators with Opus as the model being checked (the rerun): the audit's labels stay the reference, shares are of Opus's 5,595 posts, and where the audit didn't sample, Opus's labels fill in, with the count stated.

It tested ten of the claims I'd published: 4 hold, 3 hold with a correction and 3 fail. The failures are why this page no longer sorts posts into buckets. "Cost-only" (23.3% of posts) was the bin every other measured post fell into, and its description fits about 12.5% (9.8 to 16.6). And not one of the 91 "materially different" builds did something that was unavailable before. The other two fails, the share of posts that compare Jev with nothing and the share that are meta, rest on the audit taking the labels as right wherever it didn't sample. On its own random samples, the Sonnet 5 labels' 81.4% and 21.1% sit inside the intervals (77 to 85, and 16 to 25). That's why the audited numbers on this page are ranges.

What was checked

GroupHowPosts in the groupAudited
Measured production (Sonnet label)every post1515
The 91 candidate buildsevery post9191
Voice, live chat or collaborationevery post156156
Classic ML, rules or vendor baselineevery post201201
No comparison namedrandom sample4,647300
Demo, no numbersrandom sample3,681200
Measured demorandom sample1,928200
Not a realtime family, not gamesrandom sample4,475200
Realtime flag falserandom sample3,971200
Realtime flag true, not gamesrandom sample720100
Other or metarandom sample1,203150
Under 300 ms (frame, feel or turn)random sample1,064150
Unique posts audited1,681 (422 in the census groups)

How often the audit agrees with the model

The share of audited posts where the audit's label matches the Claude Opus 5.5 label, field by field. The groups aren't a random sample of the week, so the last row is agreement on the audited posts, not on every post.

GroupPostsFamilyTierEvidenceBaselineRealtimeProduction
Measured production (Sonnet label)1587%73%80%93%100%93%
The 91 candidate builds9184%75%89%92%88%99%
Voice, live chat or collaboration15679%75%97%97%93%99%
Classic ML, rules or vendor baseline20182%69%95%79%97%99%
No comparison named30082%75%93%96%96%98%
Demo, no numbers20082%71%92%98%92%98%
Measured demo20084%73%94%93%96%100%
Not a realtime family, not games20082%80%91%96%97%98%
Realtime flag false20076%80%95%92%98%98%
Realtime flag true, not games10083%78%95%94%90%96%
Other or meta15079%83%89%95%97%99%
Under 300 ms (frame, feel or turn)15093%72%97%97%89%99%
All audited cards1,68182%75%93%94%95%98%

The Claude Sonnet 5 labels the audit sampled from agree less on every field:

FieldSonnet 5 (first labels)Opus 5.5 (this page)
Family71% (κ 0.67)82% (κ 0.79)
Tier62% (κ 0.52)75% (κ 0.69)
Evidence89% (κ 0.79)93% (κ 0.86)
Baseline89% (κ 0.73)94% (κ 0.83)
Realtime86% (κ 0.67)95% (κ 0.86)
Production98% (κ 0.52)98% (κ 0.64)

The ten claims

As first published, with the audit's corrected value and its 95% interval.

#ClaimPublishedAudit on SonnetVerdictAudit on OpusVerdict on Opus
1Comparison baselines81.5% compare Jev with nothing, 11.9% with a frontier LLM, 3.0% with a small LLM, 3.6% with the thing Jev would replace78.2% nothing, 13.5% frontier LLM, 4.8% small LLM, 3.5% the tools Jev would replace; 75.7–79.8%; 12.8–15.3%; 3.9–6.6%; 2.8–5.1%Fails80.0% (77.4–81.6%); 11.3% (10.5–13.1%); 5.0% (4.2–6.9%); 3.7% (3.0–5.4%)Holds
2Production with numbersat most 1.6%; 0.9% on a re-read; 0.3% on the strict label0.14% (8 of 15 strict cards hold); 0.14–2.0%Holds with correction0.14% (0.14–2.01%)Holds with correction
3The bucket schemeabout 38% demos with no numbers; about 12% hype; about 23% cost-only; 1.6% (90 posts) materially different, 87 outside gamesthe demo and hype rules rest on the framing field, since withdrawn; 12.5% of posts fit the cost-only description, not 23.3%; 38 of the 91 still meet the material test (0.67%; 2.0% with misses), and none did something unavailable before; 9.8–16.6% cost-only; 1.2–4.6% materialFailsno measurement 67.3% (63.5–70.5%); measured demo 32.5% (29.4–36.3%); demos, no numbers (chip rule) 46.6% (41.8–51.6%); claim, no number (chip rule) 2.5% (1.3–5.4%); cost-only as described 12.5% (9.7–16.6%); material test 2.0% (1.2–4.6%)The ladder shares hold
4Voice, live chat and collaboration2.8% of posts, none in production3.8% of posts; 0 in production; 0 claiming production; 2.7–6.3%Holds with correction4.0% (2.9–6.6%); 0 in measured productionHolds
5A live loopabout a fifth of posts; 5 to 10 percent outside games7.8% outside games; about a fifth with games (21.7%); 5.8–11.1% outside gamesHoldsoutside games 8.1% (6.0–11.4%); all 22.4%Holds
6Decisions under 300 msof posts that need a decision in under 300 ms, about 91% are games90.8% are games; the set is 9.4% of posts, not 19%; 82.2–95.5% games; 8.0–10.9% of postsHoldsset 11.0% (9.5–12.5%); games 90.8% (82.2–95.5%)Holds; the set's size does not
7Posts that never say what Jev decidesa fifth never say what Jev decides, or are benchmarks or wrappers; memes and hot takes about 1% each14.9% stay meta; 12.4% vague, a benchmark or a wrapper; memes 0.28%, hot takes 0.42%; 13.3–16.3% still metaFailsmeta 20.9% (19.4–22.2%); memes 0.29%; hot takes 0.44%Holds with correction: the fifth holds on the audit-only estimate; memes and hot takes are under half a percent each
8Attentionhalf of all views sit on 1% of poststhe top 1% (58 posts) hold 53.4% of views; 49.6% without the 9 suspect posts; 47.5% of likes; a census, no samplingHolds with correctiontop 1% (56 posts) hold 53.3% of views; 49.8% without the suspect postsHolds with correction
9Jev as a classifierJev agreed with the model on 77% of cards, and was right 27 of 28 times at 0.99 or more76.7% (4,561 of 5,943); 27 of 28; 98.1% on the audit's 464 cards at 0.99 or more; a census of Jev's callsHoldsJev = Opus 78.4%; 97.0% at 0.99Holds
10The claimed multiplesthe median claim is about 28× cheaper and 6× faster28× cheaper (126 chips), 6× faster (299 chips); 34× and 8.5× without the 1× chips; a census of the chipsHolds28x cheaper (126 chips), 6x faster (297 chips)Holds

The numbers this page uses

Each audited estimate with its 95% interval, as a share of Opus's 5,595 posts. "Samples only" uses the audit's random samples where it didn't look, instead of Opus's labels. The last column is the audit's own estimate on the Sonnet 5 labels.

Posts thatAuditedSamples onlyOpus labelsAudit on Sonnet 5
Compare Jev with nothing80.0% (77.4 to 81.6)80.7% (76.8 to 84.8)80.4%78.2% (75.7 to 79.8)
Compare it with a frontier LLM11.3% (10.5 to 13.1)10.3% (6.9 to 14.2)11.3%13.5% (12.8 to 15.3)
Compare it with a small LLM5.0% (4.2 to 6.9)5.3% (2.9 to 9.9)4.3%4.8% (3.9 to 6.6)
Compare it with the tools it would replace3.7% (3.0 to 5.4)3.7% (2.9 to 7.3)4.1%3.5% (2.8 to 5.1)
Measured nothing67.3% (63.5 to 70.5)68.5%67.2% (63.4 to 70.3)
Demo with no numbers46.6% (41.8 to 51.6)47.9%46.5% (41.6 to 51.4)
A cost, speed or accuracy claim with no number2.5% (1.3 to 5.4)2.1%2.4% (1.2 to 5.3)
Measured demo32.5% (29.4 to 36.3)30.8%32.7% (29.6 to 36.5)
Measured production0.14% (0.14 to 2.01)0.66%0.14% (0.14 to 1.99)
Voice, live chat or collaboration4.0% (2.9 to 6.6)3.9% (2.7 to 7.9)3.7%3.8% (2.7 to 6.3)
A live loop outside games8.1% (6.0 to 11.4)8.3% (5.9 to 13.3)6.7%7.8% (5.8 to 11.1)
Need a decision in under 300 ms11.0% (9.5 to 12.5)10.3% (8.2 to 14.3)14.5%9.4% (8.0 to 10.9)
Games, of those under 300 ms90.8% (82.2 to 95.5)88.1%87.4%90.8% (82.2 to 95.5)
Meta (no use case, a benchmark or a wrapper)20.9% (19.4 to 22.2)20.7% (16.9 to 25.9)22.4%14.9% (13.3 to 16.3)

What the audit couldn't check

The audit drew its groups on the Sonnet 5 labels, so for Opus's labels some groups are thin: it read only part of the posts Opus puts in voice, live chat and collaboration, for example. The 847 posts the Sonnet 5 labels say compare Jev with a frontier or small LLM weren't sampled, so those estimates take Opus's word there. The posts from those groups the audit happened to label for other reasons agree 115 of 151 and 27 of 40 with the Sonnet 5 labels, and that slice isn't a random sample. The audit re-read the 15 posts the Sonnet 5 labels called production; its first-pass labels call three more posts production, which its count leaves out (Opus calls all three production too), and the 2% upper end allows for posts like those. The first review's re-read of the first pass's 95 production posts wasn't repeated. Games were left out of the check for missed voice, live chat and collaboration posts. The 78 proposal and commentary posts stayed on the model's label. The feed cuts post text at 400 characters and a claim chip can invent a number: 48 of the 1,023 posts the audit calls demos with no numbers still carry a numeric chip. And the auditor is another AI model, not a person. I labeled 13 posts myself, enough to drop two labels but not enough to check the rest, so human labels are still the missing check.

How the posts were labeled

  • Data. OpenChamber's Jev feed, snapshot 23 September 18:47 UTC: 5,950 posts from 4,593 authors, which OpenChamber selected with its own filter for what counts as a build. That includes 6 posts an earlier snapshot held and the feed later dropped. Post times decoded from the X ids run from 16 September 00:28 to 23 September 18:44 UTC. Views are X impressions and likes are X likes, both as the feed recorded them.
  • Open data. The labels from every model and the audit, the tables behind every chart, the rubric and the code are at github.com/mattheworiordan/jev-landscape. The feed's post text isn't republished there; the repo says how to fetch it.
  • Rubric. The v2 rubric has 15 families (the fourteen on the charts, and one for posts that don't use Jev), 7 latency tiers, 5 evidence levels, 5 framings, 6 baselines, and two flags: realtime infrastructure and a production claim. The framings and the seven tiers were labeled but aren't charted; they are in the labels file. "Measured" needs a number from the author's own run; TypeSafe's launch numbers quoted as Jev's general speed or price don't count. "Unclear" is allowed and preferred to a guess.
  • Models. Claude Opus 5.5, with adaptive thinking, labeled every post through Vercel AI Gateway, in batches of 40, with the v2 rubric as a cached system prompt (data/classified-opus.jsonl). Its safety filter refused one post, which is left out. Claude Sonnet 5 made the first two passes, and its v2 labels are the ones the audit sampled from. The Gateway ignores temperature for both models, so both runs were sampled at its default and aren't deterministic: on the same 120 cards, two Sonnet runs matched on family for 102 of 120. Jev classified the family of every post again as a second classifier, one call per post.
  • Noise sub-types. A separate Claude Sonnet 5 pass sorted the noise posts (Other or meta, or commentary in another family) into sub-types, and the hot takes by stance, against report/rubric-noise.md.
  • Cost. At Gateway list prices the first Sonnet pass cost $7.71 and the second $8.76 ($8.39 for the full run and the refreshes, $0.37 for two 120-card pilots). The Claude Opus 5.5 run cost $29.26. Sorting the noise into sub-types (figure 1b) cost $1.47, and Jev's pass $0.33.
  • Base. 7 duplicate posts were merged, and the 347 posts that don't use Jev (5.8%: local clones, distillations, Jev-compatible APIs over other models, builds on other models) are left out of every use-case chart, as is the post Opus refused. That leaves 5,595 posts.
  • Agreement. The audit is the main check: Grok labeled 1,681 posts blind, and agrees with Opus on family for 82% of them, on evidence for 93% and on tier for 75%. An earlier check, 120 random posts labeled blind by Claude Opus 5.5 in a separate run, is in the report. I labeled 13 posts myself as a calibration. The readers that checked every number are AI models; a proper human sample is the check still missing.

What the charts count

  • Figure 1. No measurement: a demo with no numbers, a claim with no number, a proposal or commentary. Measured demo: a number from the author's own run. Production: a number from a live system. The audit re-read every post the Sonnet 5 labels called production and 8 hold. Opus labels those 8 production too, but of all the posts it labels production the audit read 20 and agrees with only 11, so the count here is the audit's. The other seven posts the audit re-read are tests, a figure measured in development, or a claim with no number from the live system. The first labeling pass said 1.6% (95 posts), and its own re-read kept 51, a looser reading the audit did not repeat.
  • Figure 1b. Noise is a post in Other or meta, or commentary in another family. Unrelated or unclear: not about Jev, or the author's own app or demo where the post doesn't say what Jev decides. Benchmarks: tests of Jev itself with no application. Tooling: SDKs, clients, ports and infrastructure for calling Jev. On the audit's re-read memes are 0.29% of posts and hot takes 0.44%, and of the 18 hot takes 6 are bullish, 7 skeptical, 4 mixed and 1 neutral. The label moves both ways: the audit re-read 150 of the posts the Sonnet 5 labels call meta and moved 44 of them to a real use, most often classification, and it also found meta posts among the ones the labels gave a use. When I labeled posts myself, I gave a real use to all four that both models call meta, so a careful human reader would likely put the share lower.
  • Figure 1c. The 91 are every post the Sonnet 5 labels called a measured decision inside a live system, made while a person waits (100 ms to 1 s), and compared with nothing, rules, classic ML or a frontier LLM. On the audit's own labels 38 of them still meet that test, and 6 are a second post about a build already counted. The largest jobs: feed or comment filters (22), typing or form fill (13) and voice commands (13).
  • Figure 2. Views are X impressions and likes are X likes, as the feed recorded them, unverified. The top post is 7.8% of views at a like rate of 0.06%, against a median of 0.76% for posts with 1,000 or more views; without it the top 1% hold 49.7%. Posts from 16 to 19 September are 50% of posts and hold 79% of views: older posts had longer to collect them. The audit recounted these on the 5,709 posts it sampled from and got 53.4% and 49.6%.
  • Figure 3. A comparison counts only when it is named or clearly implied, such as "replaced my keyword filter" (rules) or a before-and-after figure. The frontier and small-LLM groups weren't sampled, so their estimates rest on Opus's labels there. All four of Opus's shares are inside the audit's intervals.
  • Figure 4. Bars are Opus's evidence labels by family, sorted by the share that measured anything. The darkest segment is Opus's production label, which the audit doesn't support: of the posts it labels production, the audit read 20 and agrees with 11.
  • Figure 5. The feed extracts the claim chips from the full post. The "1×" chips (8 cost, 51 speed) are an extractor artifact; without them the medians are 34× and 8.1×. Accuracy and latency chips aren't charted: they mix Jev's numbers with the baseline's and with Jev's stated confidence.
  • Figure 6. Under 300 ms is the frame, feel or turn tier. The rubric puts ordinary game input in the feel tier, so part of the games share is built in. Opus's own count, 14.5%, is outside the audit's interval. Jev's p95 call took 654 ms.
  • Figure 7. Bars are Opus's realtime flags. Sonnet 5's flags were generous (the audit kept 51 of the 100 it re-read); Opus's 6.7% sits inside the audit's interval. Opus labels two posts in this slice measured production: the audit re-read one, a live news feed, and it didn't hold; it never saw the other, a launch-alert model for trading. Voice, live chat and collaboration are 4.0% of posts on the audit (2.9 to 6.6), with none measured in production and two claiming it without a number, a live news feed and a Discord moderation bot. With games, a live loop is about a fifth of all posts (22.4%).
  • Figure 8. Views favor older posts. The posts Opus labels measured production are left out: the audit confirms 8 production posts in all. This chart replaces one of views by what a post leads with, the framing field that is dropped.
  • Figure 9. Each column is one tenth of Jev's 5,949 calls, ranked by the probability Jev stated for its answer; Jev chose among the first pass's 14 families. Overall it picked Opus's family for 78.4% of cards (kappa 0.75), and 38% of its calls at 0.99 or above are games or trading, the easiest families. On 120 cards labeled blind by Claude Opus 5.5 in a separate run, not by people, Jev is right on 27 of 28 of its calls at 0.99 or above.
  • Figure 10. The feed's first post is at 00:28 UTC on 16 Sep and the last at 18:44 UTC on 23 Sep, so both are partial days.

The buckets are gone

The first versions of this page put every post in one of six buckets: noise, hype, demo, cost-only, fast loop and material. The rules were finished after the data arrived, and the audit showed they didn't hold (claim 3), so the buckets are withdrawn. The measurement ladder in figure 1 and the substance test in figure 1c replace them. The bucket tables are kept for the record in the repo (report/v1/buckets-on-v2-labels.md), and so is the Sonnet version of this analysis (report/v2/).

What changed from the first pass

A critical review of the first pass found that "hype" rested on Sonnet calling any build "capability", that the accuracy and latency chip medians mixed Jev's numbers with baselines, and that "5,782 people" was 5,782 posts from 4,469 authors. The v2 rubric moved measured production from 1.6% to 0.2% of posts and production claims from 4.7% to 1.8%. The realtime flag went the other way (25.8% to 29.8%) because v2 flags nearly every game; figure 7 corrects for that with the audit.

Data quality and limits

  • Post text in the feed is capped at 400 characters, and 983 posts (17%) are cut. The claim chips come from the full post, so some numbers are visible only as chips.
  • The most-viewed post has a like rate of 0.06%, and 9 posts with 100,000 or more views and a like rate under 0.2% hold 12.4% of views. Figure 2 gives the numbers without them.
  • Every label is one model's reading of a short post, not a check of what was built. Claims on the cards are the authors' own and aren't reproduced here.
  • Every table behind these charts, the rubric and the scripts are in the repo; this page and its charts are generated from those CSVs by report/site/scripts/charts.mjs.