THE COMMUNITY RESOURCE
DIRECTORY
Small decisions.
Small decisions.
Endless possibilities.
Find what people build with Jev. Explore tools, libraries, experiments and ideas from across the TypeSafe ecosystem.
11,066unique resources
29clear categories
220focused subcategories
Yourssave a personal reading list
YOUR NEXT BUILD STARTS HERE
Download catalog
Explore the collection
Browse groups
Benchmarks & evaluation917 of 917 matches
Accuracy & comparisons86 of 86 matches
| Resource | Stars | Description | Category | Save |
|---|---|---|---|---|
| agentic-jevgithub.com | 1 | ⚡ AgenticJev · 拾意 — 推荐智能体实验室。搜索 × 推荐 × 广告,Jev / MiniCPM / GPT / Claude 多模型决策与生成式组合对照。 | Benchmarks & evaluationAccuracy & comparisons | |
| AIHOT pre-filter comparisonx.com | — | Benchmark of an is-this-AI-related pre-filter for AIHOT: Jev at a 30% threshold scored 98.91% versus 100% for GLM 5.3 Flash, and cost less than DeepSeek V4.1 Flash but twice Qwen3.7 Flash. | Benchmarks & evaluationAccuracy & comparisons | |
| anessbelbati/jev-rerank-benchgithub.com | 9 | Can a decision model beat dedicated rerankers? TypeSafe Jev vs Cohere Rerank 4 vs ZeroEntropy zerank-2 vs a chat-model baseline: 14 datasets, every raw API response, bootstrap ranges on every gap. | Benchmarks & evaluationAccuracy & comparisons | |
| battleship-vs-jevgithub.com | 0 | A 60-game benchmark of TypeSafe's Jev evaluation model playing Battleship. The model matches plain code; it does not beat it. | Benchmarks & evaluationAccuracy & comparisons | |
| Common Test AI comparisongithub.com | 1 | Static dashboard comparing Jev with OpenAI luna, terra, and sol variants on 836 questions across 23 subjects of Japan's 2026 university entrance Common Test, including speed and cost. | Benchmarks & evaluationAccuracy & comparisons | |
| Common Test AI comparison — appjev-luna-kyotsu-bench.shibadogcap.com | — | Static dashboard comparing Jev with OpenAI luna, terra, and sol variants on 836 questions across 23 subjects of Japan's 2026 university entrance Common Test, including speed and cost. | Benchmarks & evaluationAccuracy & comparisons | |
| Industrial email classification benchmarkx.com | — | Benchmark on 1,565 German and English supplier emails in 10 categories: Jev scored 96.4% vs Gemini's 97.5% and 98.5%, at $0.08 per 1,000 emails, and none of its 737 answers at 99%+ confidence were wrong. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev as a judgegithub.com | 97 | Experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra and Claude Sonnet 4.6 as judges of the same fixed agent runs, measuring binary accuracy, score reliability, cost and latency. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev as a reranker (hev mind)x.com | — | Untuned Jev reranker reaches 0.501 mean nDCG@10 versus Voyage rerank-3's 0.504, with a prompt tuned only on SciFact's train split. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev as a reranker (hev mind) — articlehevmind.com | — | Untuned Jev reranker reaches 0.501 mean nDCG@10 versus Voyage rerank-3's 0.504, with a prompt tuned only on SciFact's train split. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev code review benchmarkgithub.com | 5 | Benchmark of Jev, Gemini Flash and Claude Fable checking Python code against four review rules over 360 calls each: Jev cost 45x less than Flash with a 0.75 s median, but scored 98% correctness versus 100%. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev code review benchmark — demonews.ycombinator.com | — | Benchmark of Jev, Gemini Flash and Claude Fable checking Python code against four review rules over 360 calls each: Jev cost 45x less than Flash with a 0.75 s median, but scored 98% correctness versus 100%. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev court-ruling annotationgithub.com | 0 | FGV Direito SP experiment annotating 12 variables in 120 São Paulo civil court rulings: Jev scored 96.6% accuracy at 0.32 s and $0.25 per thousand rulings, versus 98.8% for Gemini 3.8 Flash at $6.56. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev court-ruling annotation — reportlab-dados.github.io | — | FGV Direito SP experiment annotating 12 variables in 120 São Paulo civil court rulings: Jev scored 96.6% accuracy at 0.32 s and $0.25 per thousand rulings, versus 98.8% for Gemini 3.8 Flash at $6.56. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev on 100 agent tool callsnews.ycombinator.com | — | Archestra compares Jev, Sonnet 5, and open-weight models on 100 real Claude Code tool calls for its information-flow annotator, where always answering benign already scores 79%. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev on 100 agent tool calls — articlearchestra.ai | — | Archestra compares Jev, Sonnet 5, and open-weight models on 100 real Claude Code tool calls for its information-flow annotator, where always answering benign already scores 79%. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev on NASA Kepler signalswww.reddit.com | — | Test of Jev on 8,054 Kepler Objects of Interest, choosing confirmed planet, false positive, or candidate from 21 measurements; 54.2% vs a 64.4% rule baseline, 72.5% after reformatting inputs. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev on NASA Kepler signals — datagist.github.com | — | Test of Jev on 8,054 Kepler Objects of Interest, choosing confirmed planet, false positive, or candidate from 21 measurements; 54.2% vs a 64.4% rule baseline, 72.5% after reformatting inputs. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev on real-world search tasks at Parallelx.com | — | Parallel's test of Jev on search reranking and related tasks, where zero-shot Jev matched at least one of their fine-tuned internal rerankers on NDCG@10. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev on ThaiExamgithub.com | 0 | Chart pages comparing jev-1.13.0 with 110 other models on Thai standardized exams: 70.7% accuracy on 567 questions at 0.35 s and $0.000029 per question, with 7.5 pp calibration error. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev RAG Benchmarkhuggingface.co | — | Frozen-candidate RAG evaluation of Jev 1.13 as reranker against OpenJev and an NVIDIA cross-encoder: on SciFact Jev reaches 79.29% nDCG@10 vs 78.70%, at about 4 s vs 307 ms p50 latency. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev RAG benchmark leaderboardhuggingface.co | — | Static leaderboard of real runs comparing Jev 1.13, OpenJev and an NVIDIA cross-encoder as rerankers on English XQuAD and SciFact RAG tasks. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev RAG Benchmark — projectgithub.com | 1 | Frozen-candidate RAG evaluation of Jev 1.13 as reranker against OpenJev and an NVIDIA cross-encoder: on SciFact Jev reaches 79.29% nDCG@10 vs 78.70%, at about 4 s vs 307 ms p50 latency. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev reranking benchmarkx.com | — | Benchmark of Jev as a search reranker on 33,047 Agent Skills Hub entries: alone it lifted NDCG@10 by only 0.012 over bge-m3, while RRF fusion of both reached 0.864. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev retrieval on QASPERx.com | — | Retrieval test on 34 QASPER questions where Jev beat pgvector with OpenAI text-embedding-3-small 17-3 (14 ties) on gold evidence coverage. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev skill suggestion on real transcriptsgithub.com | — | Measurement of Jev picking the right skill per turn over 407 skills and 356 turns mined from 838 Claude Code transcripts: 73.3% wrong loads vs 96.5% for a keyword baseline, with first-call recall identified as the ceiling. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev skill suggestion on real transcripts — repogithub.com | — | Measurement of Jev picking the right skill per turn over 407 skills and 356 turns mined from 838 Claude Code transcripts: 73.3% wrong loads vs 96.5% for a keyword baseline, with first-call recall identified as the ceiling. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs BERT on Kagglex.com | — | Kaggle experiment finding zero-training Jev slightly below a fine-tuned BERT but on par with Fable and Astra and ahead of TF-IDF logistic regression, with Noul plus a tuned threshold scoring best. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs Claude, GPT-6, Kimi, MiniMax and DeepSeekmedium.com | — | Runs the same 200 typed decisions through Jev and five frontier LLMs, comparing accuracy and cost and checking how each handles cases where human raters disagree. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs fine-tuned Qwen classifierx.com | — | Field note comparing zero-shot Jev with a fine-tuned Qwen classifier on an internal benchmark: within ~5 points of recall at matched precision, for about $70/mo at full volume. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs Gemini Flash Lite news taggingx.com | — | Comparison on a 1,000-article test set of People's Daily news tagged for Hubei relevance: Jev took 0.35 s per article versus 3 s for Gemini Flash Lite, disagreeing on about 15%. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs Gemini, DistilBERT and LightGBMx.com | — | Japanese classification benchmark comparing Jev with Gemini, DistilBERT and LightGBM, concluding Jev is a safe default for classification tasks; experiment code is on GitHub. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs Gemini, DistilBERT and LightGBM — articlezenn.dev | — | Japanese classification benchmark comparing Jev with Gemini, DistilBERT and LightGBM, concluding Jev is a safe default for classification tasks; experiment code is on GitHub. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs Laya head to headx.com | — | Side-by-side run of Jev against the open Laya model showing Laya is much faster while Jev makes much better judgments. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs Laya smoke testx.com | — | Quick side-by-side of Jev and the open Laya model: accuracy 0.727 for Jev, soft accuracy 0.580 vs 0.471, ECE 0.144 vs 0.213, latency ~710ms vs ~30-40ms, and ~$0.0004 vs ~$0 per decision. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs LLMs, BERT and Layax.com | — | Healthcare voice-AI team benchmarks Jev against Claude Sonnet 5, GPT-5-mini, a fine-tuned BERT and two open-weight models on 1,500 examples, finding near-frontier accuracy at 1/50th the cost. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs Lunagithub.com | 0 | Reproducible benchmark comparing Jev via OpenRouter's Decisions API with GPT-5.6 Luna on classifying reviews by topic, sentiment, stars, reply need, and product defects, with accuracy, latency, and cost. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs Mistral and Gemini for event validationnearhere.events | — | Use-case study pitting Jev against Mistral Small 4 and Gemini 3.5 Flash-Lite at rejecting unsuitable local-event listings; Jev scored 96% (48/50) at 0.59s and $0.043 per 1,000 decisions. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs Mistral and Gemini for event validation — discussionwww.reddit.com | — | Use-case study pitting Jev against Mistral Small 4 and Gemini 3.5 Flash-Lite at rejecting unsuitable local-event listings; Jev scored 96% (48/50) at 0.59s and $0.043 per 1,000 decisions. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs OpenAI on intent classificationtowardsdatascience.com | — | Compares Jev with two OpenAI models on the 77-class Banking77 intent dataset: 79.0% accuracy vs 83.9% and 86.2%, almost 2x faster, with well-calibrated confidence; cutting to 7 labels closes the gap. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs OpenAI on intent classification — codegithub.com | — | Compares Jev with two OpenAI models on the 77-class Banking77 intent dataset: 79.0% accuracy vs 83.9% and 86.2%, almost 2x faster, with well-calibrated confidence; cutting to 7 labels closes the gap. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs the BERT familyx.com | — | Chinese open experiment comparing Jev with zero-shot BERT-family classifiers on AG News, SST-2, Banking77, TweetEval, PAWS and a post-launch arXiv set, with 95% confidence intervals; Jev wins all 7 evaluation sets. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs. Terra on knowledge benchmarkswww.reddit.com | — | Compares Jev with GPT-5.6 Terra (reasoning off) on multiple-choice benchmarks such as MMLU, GPQA, WinoGrande, and HellaSwag; Jev is Terra-tier except on math. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev vs. Terra on knowledge benchmarks — sourcex.com | — | Compares Jev with GPT-5.6 Terra (reasoning off) on multiple-choice benchmarks such as MMLU, GPQA, WinoGrande, and HellaSwag; Jev is Terra-tier except on math. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-decision-benchmarksgithub.com | 1 | Independent evaluation of Jev 1.13 on MetaTool, When2Call, and BFCL V4 tool-selection and abstention tasks, with bilingual comparison tables against ChatGPT, Claude, Qwen, and others. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-demogithub.com | 0 | Independent demo of TypeSafe's Jev model: typed decisions with probabilities, measured side by side with OpenAI on support-ticket triage. Live local app plus a recorded replay page. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-evalgithub.com | — | Third-party comparison of Jev with gpt-4o-mini and Claude Sonnet 4.5 under identical conditions on routing 60 synthetic booking inquiries in 4 languages for a photo-shoot service in Japan. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-headline-benchgithub.com | 0 | Tests whether jev-1.13.0, seeing only the two headlines, can pick the winner of real Upworthy A/B tests: 64.5% on 10,984 randomized experiments, rising to 74.7% when the difference was decisive. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-lmgithub.com | 6 | Word-level language model that uses Jev as its output layer, with an n-gram drafter and Noul chunk verification; on held-out text Jev scored 6.92 bits/token against 6.18 for a unigram table. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-rerank-bench — write-upanessbelbati.com | — | Reranking benchmark of Jev against Cohere Rerank 4, ZeroEntropy zerank-2 and a chat-model baseline on 14 datasets; Jev's rubric averaged 0.692 against 0.691 for Cohere Pro, with no clear winner. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-rerankinggithub.com | 7 | Zero-shot reranking experiments comparing Jev with monoBERT and published TREC runs on MS MARCO and TREC-1 WSJ; on TREC DL 2021 documents Jev scored MAP 0.2790 and P@10 0.8930. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-routing-experimentgithub.com | 2 | Tests Jev as an LLM router on LLMRouterBench and RouterArena; a Jev-difficulty plus retrieval router scored 62.4% vs 60.3% for the best single model, but a no-Jev ablation matched it. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-snakegithub.com | 0 | Head-to-head Snake benchmark racing Jev against frontier LLMs from OpenRouter — same board seed, same 40-second clock, entirely in the browser. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-spam-evalgithub.com | 2 | Zero-shot ham/spam/phishing classification with Jev: 98.64% on 5,733 emails from written category definitions versus 98.87% for a TF-IDF model trained on ~4,600 labels per fold. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-tetris-benchmarkgithub.com | 0 | Reproducible Tetris decision benchmark comparing TypeSafe Jev with Claude Haiku | Benchmarks & evaluationAccuracy & comparisons | |
| jev-trace-classifiergithub.com | 0 | Tests whether a Jev Noul can tell agent-written from human-written pages on all 4,579 collusion.wiki pages, head-to-head with a local Qwen model; Jev reached 0.741 accuracy against a 77.6% majority baseline. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-vs-llm-stock-policygithub.com | 0 | Side-by-side showcase: traditional LLM vs TypeSafe Jev (System One) on stock order policy decisions | Benchmarks & evaluationAccuracy & comparisons | |
| Jev-vs-MLgithub.com | 13 | Benchmark of Jev 1.13.0 against 11 classical classification pipelines on eight datasets with three seeds: Jev reaches 96.3% balanced accuracy on IMDb versus 88.4%, while classical models lead on tabular data. | Benchmarks & evaluationAccuracy & comparisons | |
| Jev-vs-ML — appquicqdev.github.io | — | Benchmark of Jev 1.13.0 against 11 classical classification pipelines on eight datasets with three seeds: Jev reaches 96.3% balanced accuracy on IMDb versus 88.4%, while classical models lead on tabular data. | Benchmarks & evaluationAccuracy & comparisons | |
| jev-vs-tfidf-benchmarkgithub.com | 0 | Does a System One model actually beat keyword matching? A reproducible benchmark on catching disguised duplicate thesis titles. 24 cases, real production baseline, raw data and charts included. | Benchmarks & evaluationAccuracy & comparisons | |
| Jevalsgithub.com | 0 | (jevals.com): independent benchmark of TypeSafe's Jev vs LLMs | Benchmarks & evaluationAccuracy & comparisons | |
| JevBenchbenchmarkheaven.com | — | Benchmark of Jev-class decision models that ranks Jev, its open rebuilds, and instruction models on intelligence, calibration, speed, and cost, 25% each as a geometric mean; Jev led at 75.3 with SemIf second at 74.6. | Benchmarks & evaluationAccuracy & comparisons | |
| JevBench — demox.com | — | Benchmark of Jev-class decision models that ranks Jev, its open rebuilds, and instruction models on intelligence, calibration, speed, and cost, 25% each as a geometric mean; Jev led at 75.3 with SemIf second at 74.6. | Benchmarks & evaluationAccuracy & comparisons | |
| JevBench — discussionnews.ycombinator.com | — | Benchmark of Jev-class decision models that ranks Jev, its open rebuilds, and instruction models on intelligence, calibration, speed, and cost, 25% each as a geometric mean; Jev led at 75.3 with SemIf second at 74.6. | Benchmarks & evaluationAccuracy & comparisons | |
| JevBench — xx.com | — | Benchmark of Jev-class decision models that ranks Jev, its open rebuilds, and instruction models on intelligence, calibration, speed, and cost, 25% each as a geometric mean; Jev led at 75.3 with SemIf second at 74.6. | Benchmarks & evaluationAccuracy & comparisons | |
| laya-jev-lab — write-upx.com | — | Independent measurements of Jev versus the open-weight Laya on an M4 Max, where Jev scored 78% and Laya 57% on 40 Chinese support tickets, plus a local-first cascade matching Jev's accuracy at about 1.8x the speed. | Benchmarks & evaluationAccuracy & comparisons | |
| llm-rankers Jev experimentsgithub.com | — | Zero-shot TREC DL19/DL20 experiments using Jev as a pointwise, pairwise, setwise and listwise reranker; listwise Score over all 100 BM25 passages in one request reached nDCG@10 0.728 on DL19 at $0.0009 per query. | Benchmarks & evaluationAccuracy & comparisons | |
| Odin R&D: Jev vs open-weight Layagithub.com | — | Model evaluation: runs the same 48 hand-written choice/score/noul rows through Jev's hosted /api/v1/systemone and a pinned open-weight Laya on a Mac (MLX, checked row-by-row against Laya's own reference code) under pre-registered refutation criteria, publishing the raw results record, 46/48 vs 39/48 accuracy, and a reproduce command. | Benchmarks & evaluationAccuracy & comparisons | |
| openjev-sglang — demox.com | — | Compatible endpoint that serves an open mixture-of-experts model on SGLang, with BoolQ and MMLU-Pro comparisons. | Benchmarks & evaluationAccuracy & comparisons | |
| polsci-open-benchgithub.com | 11 | Benchmark of local and commercial LLMs on 33 political science classification tasks where Jev 1.13 scores a mean F1 of 0.661 versus 0.714 for Claude Opus 5, at $0.036 per 1,000 items and 0.27 s median latency. | Benchmarks & evaluationAccuracy & comparisons | |
| Putting Jev through the gauntletx.com | — | Test of Jev across five areas and 25 subareas from elementary to PhD level: Jev scored 76% against DeepSeek v4.1 Flash's 93% at roughly 50 times lower cost, strongest at text classification. | Benchmarks & evaluationAccuracy & comparisons | |
| Putting Jev through the gauntlet — blogessays.brandoncarl.com | — | Test of Jev across five areas and 25 subareas from elementary to PhD level: Jev scored 76% against DeepSeek v4.1 Flash's 93% at roughly 50 times lower cost, strongest at text classification. | Benchmarks & evaluationAccuracy & comparisons | |
| RAG agent routing benchmarkx.com | — | Chinese benchmark of 100 multi-turn dialogs on whether an agent picks the right source (knowledge base, docs, web, tools, ask user): Gemma 4 31B 77.0%, Jev 61.4%, djev-spark 32.2%, SemIf 24.0%, Laya 0%. | Benchmarks & evaluationAccuracy & comparisons | |
| sgrep Jev rerank benchmarkgithub.com | — | Isolated benchmark in sgrep, a local semantic search tool for codebases and coding-agent history, that routes Jev through its reranking boundary and compares it with local Jina, ColBERT and other rerankers on a pinned dspy-go index. | Benchmarks & evaluationAccuracy & comparisons | |
| sgrep Jev rerank benchmark — repogithub.com | — | Isolated benchmark in sgrep, a local semantic search tool for codebases and coding-agent history, that routes Jev through its reranking boundary and compares it with local Jina, ColBERT and other rerankers on a pinned dspy-go index. | Benchmarks & evaluationAccuracy & comparisons | |
| smoking-extraction-benchmarkgithub.com | — | Paired benchmark on 1,000 synthetic outpatient notes comparing Jev 1.13.0 with OpenAI structured outputs on ten smoking-history fields: 92.4% vs 98.7% complete extraction, plus cost and latency. | Benchmarks & evaluationAccuracy & comparisons | |
| sysone-benchgithub.com | 6 | First independent head-to-head benchmark of System One decision models (Laya vs Jev) on byte-identical inputs | Benchmarks & evaluationAccuracy & comparisons | |
| system-one-decision-labgithub.com | 0 | Compare System One decision models (Laya & Jev) on the same evidence: one contract, per-engine confidence gates, every decision stored with the engine that made it. | Benchmarks & evaluationAccuracy & comparisons | |
| system-one-pokergithub.com | 0 | How good is Jev AI at Texas Hold'em? TypeSafe AI's System One model plays poker through Vercel AI Gateway or OpenRouter, measured against expected value. Built with the pico framework. | Benchmarks & evaluationAccuracy & comparisons | |
| Testing Jev on public and private dataamankumar.ai | — | Measures Jev over 16,000 calls against gpt-5.4-mini and gpt-5.6-luna on four public datasets and a few thousand real pipeline decisions, showing where it wins, where it breaks and how to set thresholds. | Benchmarks & evaluationAccuracy & comparisons | |
| typed-decision-benchgithub.com | 0 | Bench of typed decision models: Jev vs OpenJev vs Laya, small local LLMs and cheap hosted LLMs on the same zero-shot classification tasks, in English and French. | Benchmarks & evaluationAccuracy & comparisons | |
| TypeSafe Jev played chessdev.to | — | Runs Jev through the LLM Chess benchmark with legal moves as a Choice, landing around #59 at Elo ~243 next to mid-pack reasoning models for ~$0.0015 per game. | Benchmarks & evaluationAccuracy & comparisons | |
| TypeSafe Jev vs Claude Code: 4 models, 2 real jobsprimeline.cc | — | Pre-registered test of Jev, GPT-5.6, Opus 5 and Haiku 4.5 on two real Claude Code jobs, where the ranking flips between the jobs and the author explains why. | Benchmarks & evaluationAccuracy & comparisons | |
| TypeSafe Jev vs Claude Code: 4 models, 2 real jobs — demox.com | — | Pre-registered test of Jev, GPT-5.6, Opus 5 and Haiku 4.5 on two real Claude Code jobs, where the ranking flips between the jobs and the author explains why. | Benchmarks & evaluationAccuracy & comparisons | |
| We tested Jev on search reranking and classificationparallel.ai | — | Search API company tests Jev zero-shot on reranking, topic classification and query freshness: it matched a custom reranker at NDCG@10 of 0.7 but trailed specialized internal classifiers. | Benchmarks & evaluationAccuracy & comparisons | |
| WebJevgithub.com | — | Specialist decision model: Apache-2.0 open-weight Qwen3.5-35B-A3B fine-tune for browser agents, served by vLLM behind the same POST /v1/systemone and /api/alpha/decisions routes so a Jev client switches by changing only the base URL and key; inside the unchanged jev-ultrafast agent it completes 38.52% of 125 hand-picked real-website tasks graded by deterministic verifiers, against 16.67% for Jev 1.13. | Benchmarks & evaluationAccuracy & comparisons |
Latency & cost7 of 7 matches
| Resource | Stars | Description | Category | Save |
|---|---|---|---|---|
| 500 real-time agents in 3Dx.com | — | Benchmark running 500 real-time agents in parallel in a 3D environment, reporting 500ms average latency and 35 API calls/s with no optimizations. | Benchmarks & evaluationLatency & cost | |
| endpointgithub.com | 23 | goodrahstar/jev-column-race — Jev vs Gemini 3.8 Flash: labelling 1,000 app reviews, 4.1× faster and 7× cheaper · endpoint · JavaScript | Benchmarks & evaluationLatency & cost | |
| Jev-is-oddgithub.com | 0 | Ask Jev by TypeSafe AI whether a number is odd. TypeScript, real token usage, and latency benchmarks. | Benchmarks & evaluationLatency & cost | |
| jev-measuredgithub.com | — | Reproducible measurements of cost, latency and raw output from the live Jev API via OpenRouter across eight use cases, plus a small head-to-head on 27 support tickets; the whole run cost under one cent. | Benchmarks & evaluationLatency & cost | |
| PDF Racegithub.com | 6 | Race of three document pipelines on 12 arXiv papers: Docling with Jev, Docling with Gemini 3.8 Flash, and Gemini reading the PDF; all scored 12/12, but the Jev lane cost $0.0022 versus $0.0882. | Benchmarks & evaluationLatency & cost | |
| PDF Race — apppdf-race.vercel.app | — | Race of three document pipelines on 12 arXiv papers: Docling with Jev, Docling with Gemini 3.8 Flash, and Gemini reading the PDF; all scored 12/12, but the Jev lane cost $0.0022 versus $0.0882. | Benchmarks & evaluationLatency & cost | |
| Typed judgments or agentic loops?blog.r6i.it | — | Hierarchical Choices with fan-out against a GPT tool-calling agent, cutting average latency from 9.62 s to 1.38 s. | Benchmarks & evaluationLatency & cost |
Calibration & consistency22 of 22 matches
| Resource | Stars | Description | Category | Save |
|---|---|---|---|---|
| 64 tiny benchmarks for Jevwww.ramonov.com | — | Informal eval that sends 64 off-label questions to jev-1.13.0 50 times each (3,200 calls) and charts how stable the returned distributions are, including where the one-shot answer flips. | Benchmarks & evaluationCalibration & consistency | |
| 64 tiny benchmarks for Jev — discussionnews.ycombinator.com | — | Informal eval that sends 64 off-label questions to jev-1.13.0 50 times each (3,200 calls) and charts how stable the returned distributions are, including where the one-shot answer flips. | Benchmarks & evaluationCalibration & consistency | |
| ai_sdkgithub.com | 4 | scienthoon/jev-ood-calibration — Independent calibration test of TypeSafe's Jev on a task it cannot have seen: 900 rule-generated support tickets (choice / score / boolean) plus 3 public benchmarks via Vercel AI Gateway. Raw responses, ECE with noise floor, temperature refit, per-type sign of miscalibration. Reproducible for ~$0.06. · ai_sdk · Python | Benchmarks & evaluationCalibration & consistency | |
| ASSAY-001github.com | — | Pre-registered check of Jev's calibration and type-safety claims: calibrated on CLINC150 (ECE 0.0204), overconfident on Banking77 (ECE 0.0936), and zero type errors across 8,576 responses, with full logs. | Benchmarks & evaluationCalibration & consistency | |
| ASSAY-001: Jev calibrationdonttrustme.ai | — | Pre-registered check of Jev's calibration and type-safety claims on Banking77 and CLINC150: calibrated on CLINC150 (ECE 0.0204) but overconfident on Banking77 (ECE 0.0936), with zero type errors in 8,576 responses. | Benchmarks & evaluationCalibration & consistency | |
| calibrantgithub.com | 0 | Does your model's confidence mean anything on your data? Calibration layer for typed probabilistic decisions — reliability, ECE, Brier, recalibration maps and cost-aware thresholds from decisions + outcomes. | Benchmarks & evaluationCalibration & consistency | |
| Calibrating Jevx.com | — | Economist's calibration study across five labeled tasks (about 37,000 items, 24 phrasings each), finding Jev's probabilities often overconfident and improved by an online Foster-Hart correction. | Benchmarks & evaluationCalibration & consistency | |
| endpointgithub.com | 10 | abhixhek/jevcal — Stop guessing confidence thresholds: calibrate, threshold, and drift-check typed decision models (TypeSafe Jev) against an LLM teacher. · endpoint · Python | Benchmarks & evaluationCalibration & consistency | |
| Hosted Jev calibration studygithub.com | — | Reproducible calibration study of hosted jev-1.13.0 on 240 seeded questions (92.2% accuracy, Brier 0.048, ECE 0.041 on Noul items), plus a 200-question adversarial follow-up with ECE 0.012. | Benchmarks & evaluationCalibration & consistency | |
| Hosted Jev calibration study — adversarialgithub.com | — | Reproducible calibration study of hosted jev-1.13.0 on 240 seeded questions (92.2% accuracy, Brier 0.048, ECE 0.041 on Noul items), plus a 200-question adversarial follow-up with ECE 0.012. | Benchmarks & evaluationCalibration & consistency | |
| Hosted Jev calibration study — codegithub.com | — | Reproducible calibration study of hosted jev-1.13.0 on 240 seeded questions (92.2% accuracy, Brier 0.048, ECE 0.041 on Noul items), plus a 200-question adversarial follow-up with ECE 0.012. | Benchmarks & evaluationCalibration & consistency | |
| Hosted Jev calibration study — repogithub.com | — | Reproducible calibration study of hosted jev-1.13.0 on 240 seeded questions (92.2% accuracy, Brier 0.048, ECE 0.041 on Noul items), plus a 200-question adversarial follow-up with ECE 0.012. | Benchmarks & evaluationCalibration & consistency | |
| Jev calibration statisticshuggingface.co | — | Confidence statistics from a retrieval benchmark judged by jev-1.13.0: its scores separate answer-bearing passages (AUROC 0.899), yet the Jev-judged pipeline scored 0.612 vs 0.740 with no judge. | Benchmarks & evaluationCalibration & consistency | |
| jev-benchmarksgithub.com | 21 | Probability-aware benchmark comparing Jev with GLiNER2.5 on zero-shot text classification, measuring calibration, coverage at a fixed error budget and latency on three BTZSC datasets. | Benchmarks & evaluationCalibration & consistency | |
| jev-calibration-auditgithub.com | 0 | Independent API-only calibration audit of Jev in seven experiments (~7,000 calls): removing the abstain option drops accuracy on unanswerable items from 0.950 to 0.000, while Korean leaves calibration unchanged. | Benchmarks & evaluationCalibration & consistency | |
| jev-certifygithub.com | 1 | Finite-sample guarantees for Jev (TypeSafe's System One). Conformal risk control turns calibrated probabilities into certified routing thresholds; prediction-powered inference audits them. 2,412 decisions on CLINC150 for $0.23 — including the shift and prevalence cases where the guarantee breaks. | Benchmarks & evaluationCalibration & consistency | |
| jev-explorationgithub.com | 3 | Evidence ledger of Jev claims that recomputes public calibration results against sample-size noise and runs an 800-item difficulty gradient, finding the probabilities are not calibrated at any difficulty. | Benchmarks & evaluationCalibration & consistency | |
| jev-takes-mauboussingithub.com | 0 | Evaluating TypeSafe's Jev on Michael Mauboussin's 50-question decision calibration test | Benchmarks & evaluationCalibration & consistency | |
| jevbenchgithub.com | 1 | Preregistered, bias-corrected test of TypeSafe Jev's calibration under human disagreement (ChaosNLI, 100 labels per item) | Benchmarks & evaluationCalibration & consistency | |
| scienthoon/jev-ood-calibrationgithub.com | 4 | Independent calibration test of TypeSafe's Jev on a task it cannot have seen: 900 rule-generated support tickets (choice / score / boolean) plus 3 public benchmarks via Vercel AI Gateway. Raw responses, ECE with noise floor, temperature refit, per-type sign of miscalibration. Reproducible for ~$0.06. | Benchmarks & evaluationCalibration & consistency | |
| semantic-microscopegithub.com | 0 | Label every sentence of a document with calibrated probabilities from Jev, rendered as a heatmap | Benchmarks & evaluationCalibration & consistency | |
| Sys1Cal-v1github.com | 0 | A Benchmark for Calibration and Semantic Evaluation of System One Models | Benchmarks & evaluationCalibration & consistency |
Security evaluations9 of 9 matches
| Resource | Stars | Description | Category | Save |
|---|---|---|---|---|
| Jev jailbreak benchmarkx.com | — | Prompt-injection benchmark pitting Jev against standard guardrail classifiers including Meta's: it wins an open suite and the newest attack set, loses on older sets and cannot hold a tight false-alarm budget. | Benchmarks & evaluationSecurity evaluations | |
| Jev jailbreak benchmark — articlebacknotprop.com | — | Prompt-injection benchmark pitting Jev against standard guardrail classifiers including Meta's: it wins an open suite and the newest attack set, loses on older sets and cannot hold a tight false-alarm budget. | Benchmarks & evaluationSecurity evaluations | |
| Jev jailbreak benchmark — discussionnews.ycombinator.com | — | Pits one Noul per message against four local injection detectors, finding strong ranking but little recall at a 1% false-positive rate. | Benchmarks & evaluationSecurity evaluations | |
| jev-labgithub.com | 0 | Real browser-agent safety evaluation: Jev versus a baseline on benign and injected tasks | Benchmarks & evaluationSecurity evaluations | |
| jev-sec-benchgithub.com | 3 | Blind security benchmarks for jev-1.13.0 with a Go runner and TUI: 96.5% accuracy on all 662 deepset prompt-injection messages at a plain 0.50 cut, plus 200 vulnerable-code pairs. | Benchmarks & evaluationSecurity evaluations | |
| jev-sec-benchx.com | — | Blind security benchmarks for Jev, TypeSafe's System One model: prompt injection and vulnerable code detection, built on jev-go | Benchmarks & evaluationSecurity evaluations | |
| kiarina/labs: safety judgmentgithub.com | — | moderation plus shell-command safety checks. Unique: strong non-English (Japanese) moderation: 36 misses vs OpenAI's 292 on 826 harmful texts. [self-reported] | Benchmarks & evaluationSecurity evaluations | |
| primitivesgithub.com | 3 | Gaurav-Gosain/jev-sec-bench — Blind security benchmarks for Jev, TypeSafe's System One model: prompt injection and vulnerable code detection, built on jev-go · primitives · Go | Benchmarks & evaluationSecurity evaluations | |
| typesafe-guardrailsgithub.com | 1 | Rebuilding the $1 Chevy Tahoe jailbreak, then stopping it: TypeSafe System One as a guardrail in an OpenAI Agents SDK agent, traced with Arize AX | Benchmarks & evaluationSecurity evaluations |
Evaluation harnesses39 of 39 matches
| Resource | Stars | Description | Category | Save |
|---|---|---|---|---|
| auto-research System Onegithub.com | — | Research module that implements the Choice/Score/Noul contract and runs Banking77 and public-suite evaluations of Jev against a local calibrated scorer and the NanoJev, Nimble and Laya checkpoints. | Benchmarks & evaluationEvaluation harnesses | |
| auto-research System One — docsgithub.com | — | Research module that implements the Choice/Score/Noul contract and runs Banking77 and public-suite evaluations of Jev against a local calibrated scorer and the NanoJev, Nimble and Laya checkpoints. | Benchmarks & evaluationEvaluation harnesses | |
| auto-research System One — repogithub.com | — | Research module that implements the Choice/Score/Noul contract and runs Banking77 and public-suite evaluations of Jev against a local calibrated scorer and the NanoJev, Nimble and Laya checkpoints. | Benchmarks & evaluationEvaluation harnesses | |
| cultivar TypeSafe grader — repogithub.com | 41 | Optional grading backend in Pinecone's agent-skill testing CLI that scores sandboxed agent runs against task criteria with Jev instead of Claude, reported as about 30x cheaper and aimed at CI gates. | Benchmarks & evaluationEvaluation harnesses | |
| DeepEval TypeSafe integrationdeepeval.com | — | Experimental DeepEval integration routes supported metric verdicts, scores, and classifier labels to Jev. | Benchmarks & evaluationEvaluation harnesses | |
| endpointgithub.com | 1 | 4esv/jev-eval — Benchmark TypeSafe Jev against any OpenRouter model on your own labelled classification data: accuracy, calibration, latency, cost · endpoint · Python | Benchmarks & evaluationEvaluation harnesses | |
| endpointgithub.com | 1 | ElshinQ/jevaluate — Jevaluate: evaluate before you trust. Field notes, runnable scripts and an agent skill for TypeSafe Jev: gated evals, a browser loop, a product walk with DeepSeek vision, a UI text judge and a first-click tree test. Co-authored with Claude Fable 5.1. · endpoint · JavaScript | Benchmarks & evaluationEvaluation harnesses | |
| endpointgithub.com | 5 | Nainish-Rai/jev-frontend-qa — Evidence-driven frontend QA built on Jev Ultrafast and Browser Harness, with a synthetic todo demo. · endpoint · Python | Benchmarks & evaluationEvaluation harnesses | |
| endpointgithub.com | 1 | amr05008/jev-sandbox — Test bench for TypeSafe's Jev · endpoint · TypeScript | Benchmarks & evaluationEvaluation harnesses | |
| endpointgithub.com | 24 | myc0576/SmartMoney-Cub — Read-only trading journal and review harness: Jev typed judgments, agent integration, and a reproducible finance benchmark. No orders, no advice. · endpoint · Python | Benchmarks & evaluationEvaluation harnesses | |
| jev-benchmarks (thijmenkam)github.com | — | Reproducible harness that asks Jev and frontier LLMs identical typed questions about the same state and scores accuracy, calibration, consistency, latency, cost, and schema validity, with a 150-call-per-model run. | Benchmarks & evaluationEvaluation harnesses | |
| jev-decision-benchgithub.com | 2 | A reproducible, self-hosted benchmark for comparing JEV and LLM decision-making on your own datasets, models, and evaluation policies. | Benchmarks & evaluationEvaluation harnesses | |
| jev-demogithub.com | 0 | JEV Studio is a small Next.js evaluation lab for turning natural-language input into structured signals with the Typesafe SystemOne API. | Benchmarks & evaluationEvaluation harnesses | |
| jev-evalgithub.com | 1 | Harness that benchmarks Jev against any OpenRouter model or local checkpoint on labeled classification data for accuracy, calibration, latency, and cost, with results versus GPT-5.6 Terra, open-jev, Kev-0.8B, and Laya. | Benchmarks & evaluationEvaluation harnesses | |
| jev-evalsgithub.com | 1 | Rubric-based eval harness for LLM and agent outputs that sends input, output and expected once as state and answers every rubric as a Jev question in one call, cheap enough to run on every PR. | Benchmarks & evaluationEvaluation harnesses | |
| jev-harnessgithub.com | 15 | TypeScript library that turns Jev answers into shippable actions with a policy map, confidence gate, shadow mode, recipes and an offline eval CLI; a row-filter job took 1.3 s versus 48.9 s with Claude CLI. | Benchmarks & evaluationEvaluation harnesses | |
| Jev-LLM-Playgroundgithub.com | 0 | Independent playground for TypeSafe AI Jev decision models: typed decisions, support-ticket routing, reproducible evaluations, and a local browser demo. | Benchmarks & evaluationEvaluation harnesses | |
| jev-probegithub.com | 0 | Testing environment for TypeSafe's Jev model. It applies controlled perturbations (typos, adversarial text, reordering) to test cases, executes real API calls, and logs the raw JSON responses to DuckDB. Includes a web dashboard to track probability shifts, latency, and cost against an LLM baseline. | Benchmarks & evaluationEvaluation harnesses | |
| jev-quiz-pilotgithub.com | 0 | Put Jev in the pilot's seat of a web quiz. A Python CLI that lets TypeSafe's Jev navigate quizzes in your Chrome and logs every pick, so you can measure how it does. | Benchmarks & evaluationEvaluation harnesses | |
| jev-replay-labgithub.com | 0 | Local Jev evaluation workbench: datasets, typed questions, threshold simulation and run comparison | Benchmarks & evaluationEvaluation harnesses | |
| jev-research-evalgithub.com | 2 | Reproducible eval harness and field note for research-browser tasks run with jev-ultrafast, with 11 baseline cases, human and quant stress suites, QC grades, a suite runner and a report generator. | Benchmarks & evaluationEvaluation harnesses | |
| jev-sandboxgithub.com | 1 | Test bench for TypeSafe's Jev | Benchmarks & evaluationEvaluation harnesses | |
| jevalgithub.com | 0 | open-source evaluations for AI outputs and agents, judged by Jev | Benchmarks & evaluationEvaluation harnesses | |
| jevalsgithub.com | 1 | Local browser workbench for authoring Jev Noul, Choice and Score questions with example cases and expected answers, running them and comparing saved results. | Benchmarks & evaluationEvaluation harnesses | |
| jevalsgithub.com | 99 | Framework-agnostic evals and guardrails for agents that send all of a trace's checks to Jev in one request, fast enough for the agent loop, with Kev or Laya locally or a chat-model fallback. | Benchmarks & evaluationEvaluation harnesses | |
| jevcompatgithub.com | — | Conformance spec, test runner, proxy, reference mock, and GitHub Action for Jev-compatible API servers. | Benchmarks & evaluationEvaluation harnesses | |
| krinogithub.com | 0 | Krino (κρίνω) — probing, analysis, and replication toolkit for JEV-class decision models | Benchmarks & evaluationEvaluation harnesses | |
| model_idgithub.com | 97 | danielgshea/jev-as-a-judge — Using Jev as an evaluator. · model_id · Python | Benchmarks & evaluationEvaluation harnesses | |
| open-jev (DiffusionGemma)github.com | 41 | Experimental harness for typed JSON decisions with the DiffusionGemma 26B-A4B diffusion model that denoises freely and then picks the most likely allowed tokens, benchmarked on Every's TypeSafe lab tasks. | Benchmarks & evaluationEvaluation harnesses | |
| openevalsgithub.com | 4 | Affordable platform for parallel agent evals and observability. Powered by JEV. | Benchmarks & evaluationEvaluation harnesses | |
| Openwork Jev Verification Dictionarygithub.com | — | Experimental evaluator that uses Jev to select author-defined checks before deterministic replay. | Benchmarks & evaluationEvaluation harnesses | |
| sdkgithub.com | 15 | AntonioCoppe/jev-harness — Decision harness for TypeSafe Jev — confidence gates, shadow mode, recipes, and evals. Claude CLI 48.9s → Jev 1.3s on the same row-filter job. · sdk · TypeScript | Benchmarks & evaluationEvaluation harnesses | |
| sdkgithub.com | 1 | dayhaysoos/jevals — Local evaluation workbench for TypeSafe Jev · sdk · TypeScript | Benchmarks & evaluationEvaluation harnesses | |
| sdkgithub.com | 4 | memovai/openevals — Affordable for parallel online agent evals and observability. Powered by JEV. · sdk · TypeScript | Benchmarks & evaluationEvaluation harnesses | |
| sdkgithub.com | 2 | superradcompany/multiverse-of-madness — Jev and Microsandbox explore alternate game futures with a reusable TypeScript learning harness · sdk · TypeScript | Benchmarks & evaluationEvaluation harnesses | |
| sdkgithub.com | 6 | wondertwins/jev-benchmark — Benchmarks and a playground for TypeSafe's Jev (System One) model: chess, and who-is-the-player-talking-to for speech-to-text game NPCs · sdk · Python | Benchmarks & evaluationEvaluation harnesses | |
| trade-jevgithub.com | 11 | Backtest harness that tests Jev as a buy, sell, or hold trader on 15 trading days of Nasdaq futures L10 order-book data, against hold, random, and imbalance baselines, with replay and a web viewer. | Benchmarks & evaluationEvaluation harnesses | |
| typed_evalsgithub.com | 13 | Python library and CLI that evaluates LLM responses, RAG datasets, and recorded agent runs with Jev as the judge, guards tools before they execute, and can calibrate metrics against human pass/fail labels. | Benchmarks & evaluationEvaluation harnesses | |
| World Monitor Jev Headline Evaluationgithub.com | 87,060 | Experimental harness that scores news headlines against geopolitical threat levels and event categories. | Benchmarks & evaluationEvaluation harnesses |
Datasets & case studies7 of 7 matches
| Resource | Stars | Description | Category | Save |
|---|---|---|---|---|
| jev-banc-francaisgithub.com | 0 | Evaluate typed AI decisions on labeled French-language cases. | Benchmarks & evaluationDatasets & case studies | |
| jev-benchhuggingface.co | — | Human-labeled datasets reformatted into System One questions (22 configs, 166,054 rows), keeping human label distributions where they exist, with accuracy and calibration results for jev-1.13.0. | Benchmarks & evaluationDatasets & case studies | |
| jev-bench — repogithub.com | 4 | Human-labeled datasets reformatted into System One questions (22 configs, 166,054 rows), keeping human label distributions where they exist, with accuracy and calibration results for jev-1.13.0. | Benchmarks & evaluationDatasets & case studies | |
| jev-typesafe-real-financial-use-casesgithub.com | 0 | Fifty real-world financial use cases for TypeSafe's Jev model: typed, structured LLM answers over ledgers, fraud, portfolios, trades and filings, each graded against data where the right answer is known. | Benchmarks & evaluationDatasets & case studies | |
| no-hallucinationgithub.com | 3 | Three measured experiments on RAG hallucination: quote-checking, TypeSafe's Jev, and IBM's STAIR. 850+ graded questions, raw responses included. | Benchmarks & evaluationDatasets & case studies | |
| snbt-jev-benchgithub.com | 0 | Jev on Indonesia's SNBT 2025 university entrance test: 159 questions, seven subtests, audited answer keys. | Benchmarks & evaluationDatasets & case studies | |
| what-the-jevgithub.com | 4 | What can Jev actually do? Reproducible experiments and research reports exploring its capabilities and limits. | Benchmarks & evaluationDatasets & case studies |
More projects & source code643 of 643 matches
| Resource | Stars | Description | Category | Save |
|---|---|---|---|---|
| ../results/github.com | — | Results from the 2026-09-19 run are in ../results/: one JSONL line per question per model (id, gold label, the probability the model returned, end-to-end and server-side latency, input tokens, cost), plus the aggregate JSON that analyze.mjs and stats.mjs produce. | Benchmarks & evaluationMore projects & source code | |
| acoyfellow/nightglassgithub.com | 1 | Owned, deterministic classifier for checking whether agent claims are supported by evidence, with optional Jev comparison through Cloudflare AI Gateway. | Benchmarks & evaluationMore projects & source code | |
| agentic-jev — READMEgithub.com | 1 | AgenticJev is a discovery-and-shortlisting app: it retrieves candidates from sources, can use Jev or other models for selection, then incorporates feedback. Jev is an optional decision stage, not the whole search stack. | Benchmarks & evaluationMore projects & source code | |
| AI Elo Rankergithub.com | 7 | Tournament engine that ranks texts such as poems, pitches, cold emails, and ad hooks through Jev pairwise judgments, Swiss-system matchmaking, and Elo ratings, streamed to a live WebSocket dashboard. | Benchmarks & evaluationMore projects & source code | |
| ai-elo-ranker — READMEgithub.com | 7 | AI Elo Ranker uses Jev judgments in a tournament/Elo workflow to rank text entries. | Benchmarks & evaluationMore projects & source code | |
| ai_sdkgithub.com | 4 | caiovicentino/jev-align — Calibrated alignment verifier for LLM responses and agent plans — powered by Jev · ai_sdk · JavaScript | Benchmarks & evaluationMore projects & source code | |
| ai_sdkgithub.com | 11 | hellozenstrategist-lab/eutrya — Jev-native AI security harness for autonomous research, multi-agent swarms, persistent hunt boards, and long-running agent workflows. CLI-first, open source, and built for authorized security research. · ai_sdk · JavaScript | Benchmarks & evaluationMore projects & source code | |
| ai_sdkgithub.com | 1 | sstehniy/jev-calculator — iOS 6-inspired Jev calculator demo with a lifetime API budget · ai_sdk · TypeScript | Benchmarks & evaluationMore projects & source code | |
| alperenerol/jev-1.13-mini-benchmarkgithub.com | 1 | Mini benchmark of TypeSafe's jev-1.13 structured decision model (OpenRouter Decisions API) on labeled support-triage: noul/choice/score, consistency, cost, lessons learned | Benchmarks & evaluationMore projects & source code | |
| AMBER decision-axis evalsgithub.com | — | Evaluation pipeline in the AMBER replay benchmark for decision models such as Jev, reporting per-family accuracy, calibration bins, ECE, threshold sweeps and cost/latency from HMAC-signed records. | Benchmarks & evaluationMore projects & source code | |
| AMBER decision-axis evals — repogithub.com | — | Evaluation pipeline in the AMBER replay benchmark for decision models such as Jev, reporting per-family accuracy, calibration bins, ECE, threshold sweeps and cost/latency from HMAC-signed records. | Benchmarks & evaluationMore projects & source code | |
| AnyLM2Jevgithub.com | 0 | Reproducing the Jev decision interface | Benchmarks & evaluationMore projects & source code | |
| AnyLM2Jev — READMEgithub.com | 0 | Open-model research reconstructs Jev's typed decision interface on small language models and compares self-consistency-distilled readouts with raw option logits, rather than calling hosted Jev for its own decisions. | Benchmarks & evaluationMore projects & source code | |
| aotn-jev-turbine-triagegithub.com | 0 | Triage real wind turbine SCADA events with Jev typed questions, with a second look from production data for uncertain events | Benchmarks & evaluationMore projects & source code | |
| aotn-jev-turbine-triage — READMEgithub.com | 0 | Educational turbine-alarm example asks Jev yes/no questions and uses code to route wind-farm events to action, monitoring or no action; it is not production engineering advice. | Benchmarks & evaluationMore projects & source code | |
| Armature TypeSafe decision A/Bgithub.com | — | A/B benchmark inside the Armature agent harness on 20 labeled code-review statements: Jev scored 0.87 overall versus 0.98 for a qwen3.6-27b judge, at 260 ms versus 20,591 ms per call. | Benchmarks & evaluationMore projects & source code | |
| Armature TypeSafe decision A/B — repogithub.com | — | A/B benchmark inside the Armature agent harness on 20 labeled code-review statements: Jev scored 0.87 overall versus 0.98 for a qwen3.6-27b judge, at 260 ms versus 20,591 ms per call. | Benchmarks & evaluationMore projects & source code | |
| ask-jevgithub.com | 1 | PowerShell hook for Codex that runs an explicit, advisory Jev audit of a local coding session's recorded execution, judging whether claims are backed by evidence, via a :jev command. | Benchmarks & evaluationMore projects & source code | |
| ask-jev — READMEgithub.com | 1 | The Codex plugin's :jev audit and ask commands send selected conversation context to TypeSafe for Jev inference. | Benchmarks & evaluationMore projects & source code | |
| ask-twicegithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| ask-twice — READMEgithub.com | 0 | Ask Twice compares Jev and an open-source NLI model on complaint-question answers used to predict refunds; its README says TF-IDF plus logistic regression scored highest. | Benchmarks & evaluationMore projects & source code | |
| askjevgithub.com | 0 | A tree housing every closed question humans or machines ask, answered by Jev (TypeSafe System One): capabilities, defaults, and jaggedness. Not a benchmark. | Benchmarks & evaluationMore projects & source code | |
| askjev — READMEgithub.com | 0 | Question-taxonomy explorer and collection pipeline organize Jev’s typed answers to explore its judgment behavior, explicitly not as a benchmark. | Benchmarks & evaluationMore projects & source code | |
| battleship-vs-jev — READMEgithub.com | 0 | This Battleship benchmark compares Jev-driven choices with code-only strategies and visualizes per-cell probabilities. Its README emphasizes comparison methodology; this review does not validate outcome scores. | Benchmarks & evaluationMore projects & source code | |
| BizzJevgithub.com | 1 | Experiments with TypeSafe/Jev semantic gates and a Semantic Operations Lab demo. | Benchmarks & evaluationMore projects & source code | |
| BizzJev — READMEgithub.com | 1 | BizzJev is a local demo in which customer messages are sent to TypeSafe System One as separate Noul gate questions; application thresholds map answers to NO, REVIEW, or YES. | Benchmarks & evaluationMore projects & source code | |
| btc-jev-signal — READMEgithub.com | 4 | A non-trading BTC forecast experiment uses Jev with public Kraken and Binance market data; its README says it never places or prepares trades. | Benchmarks & evaluationMore projects & source code | |
| Bus 2.0 dispatch benchmarkgithub.com | 21 | Benchmark of online dispatch for an on-demand shared shuttle comparing Jev, local Laya, Claude, OpenAI and Gemini policies; its Jev notes show why arithmetic belongs in code (157.6 min² vs a 16.2 reference). | Benchmarks & evaluationMore projects & source code | |
| Bus 2.0 dispatch benchmark — notesgithub.com | 21 | Benchmark of online dispatch for an on-demand shared shuttle comparing Jev, local Laya, Claude, OpenAI and Gemini policies; its Jev notes show why arithmetic belongs in code (157.6 min² vs a 16.2 reference). | Benchmarks & evaluationMore projects & source code | |
| calibrant — READMEgithub.com | 0 | calibrant currently describes a JSONL probability-calibration engine and CLI; its TypeSafe Jev adapter is listed on the roadmap, not as a shipped feature. | Benchmarks & evaluationMore projects & source code | |
| calibre — calibregithub.com | 2 | Measure when to use Jev and other models on your data, then route accordingly. | Benchmarks & evaluationMore projects & source code | |
| calibre — READMEgithub.com | 2 | Janus routes decisions according to a measured confidence threshold and accepts typesafe:jev-latest as its primary model. | Benchmarks & evaluationMore projects & source code | |
| can-jev-bayes — READMEgithub.com | 1 | This research project studies Jev's sequential decisions under uncertainty and the role of Bayesian methods. | Benchmarks & evaluationMore projects & source code | |
| Chimera Jev governance benchmarkgithub.com | — | Preregistered benchmark in the Chimera agent repo comparing Jev (a danger Noul plus a block/review/allow Choice) with a DeepSeek judge and verbalized probabilities; Jev reached 0.903 AUROC on ambiguous items. | Benchmarks & evaluationMore projects & source code | |
| Chimera Jev governance benchmark — repogithub.com | — | Preregistered benchmark in the Chimera agent repo comparing Jev (a danger Noul plus a block/review/allow Choice) with a DeepSeek judge and verbalized probabilities; Jev reached 0.903 AUROC on ambiguous items. | Benchmarks & evaluationMore projects & source code | |
| chinese-workflow-decision-benchgithub.com | 1 | Feishu-style Chinese message classification benchmark with 64 frozen synthetic scenarios, where Jev classified 64/64 on a single Choice versus 20/64 for Laya, with latencies and raw responses published. | Benchmarks & evaluationMore projects & source code | |
| cogp-jev-lensgithub.com | 0 | COGP × JEV: AI Lens で同じ地図を見る Web GIS 実験 · TypeScript | Benchmarks & evaluationMore projects & source code | |
| cogp-jev-lens — READMEgithub.com | 0 | This experimental GIS semantically evaluates OSM tags with JEV to change how POIs appear on the map. | Benchmarks & evaluationMore projects & source code | |
| confidence and rankinggithub.com | — | For a comparison, inspect the dataset, model version, baseline settings, repeated runs, and failures. See confidence and ranking. | Benchmarks & evaluationMore projects & source code | |
| Convex decision model evalsgithub.com | 128 | Benchmark of 106 decision questions drawn from 90 Convex coding evals that compares Jev via OpenRouter's decisions API against language models, recording probabilities, confidence and cost. | Benchmarks & evaluationMore projects & source code | |
| DecaState Jev vs frontier LLMs — projectgithub.com | — | Benchmark of Jev against four frontier LLMs on 36 support-triage decisions from US-West and Singapore: about 176x cheaper and 9x faster than GPT-6 Astra, with 0 type errors. | Benchmarks & evaluationMore projects & source code | |
| DecaState Jev vs frontier LLMs — repogithub.com | — | Benchmark of Jev against four frontier LLMs on 36 support-triage decisions from US-West and Singapore: about 176x cheaper and 9x faster than GPT-6 Astra, with 0 type errors. | Benchmarks & evaluationMore projects & source code | |
| decision-model-benchmarkgithub.com | 9 | Independent, reproducible benchmark: a decision model (jev), eight constrained LLMs, and deterministic baselines on typed decisions - accuracy, calibration, latency, cost, failure modes | Benchmarks & evaluationMore projects & source code | |
| decision-model-benchmark — READMEgithub.com | 9 | An independent benchmark comparing TypeSafe Jev typed decisions with structured-output LLMs and deterministic baselines, including calibration, failures, latency and cost. Listed for its evaluation methodology and artifacts, not endorsement of a model or certification of reported rankings. | Benchmarks & evaluationMore projects & source code | |
| DecisionBridgegithub.com | 1 | Jev-inspired adapter that turns existing LLMs (OpenAI, Anthropic, OpenRouter, local MLX) into decision functions with explicit option scores, optional calibration on labeled examples and human-review thresholds. | Benchmarks & evaluationMore projects & source code | |
| decisionopsgithub.com | 1 | Jev thinks. Your code acts. The open-source lab for Choice, Score & Noul decisions. | Benchmarks & evaluationMore projects & source code | |
| decisionops — READMEgithub.com | 1 | A local DecisionOps workbench evaluates saved Jev decisions against known outcomes; the model supplies judgments while caller code retains action authority. | Benchmarks & evaluationMore projects & source code | |
| DeepSearcher Jev Stopping Evaluationgithub.com | — | Evaluation of Jev and DeepSeek stopping strategies for multi-hop question answering. | Benchmarks & evaluationMore projects & source code | |
| demo_trend_searchergithub.com | 0 | demo trend seracher using jev | Benchmarks & evaluationMore projects & source code | |
| demo_trend_searcher — READMEgithub.com | 0 | The arXiv trend-search pipeline assigns paper decisions to Jev and summaries to GPT. | Benchmarks & evaluationMore projects & source code | |
| DoomSatgithub.com | 2 | F´ flight software → CCSDS/Yamcs → Open MCT, with jev (System One) and Claude Sonnet 5 (System Two) driving Doom over that real mission stack. | Benchmarks & evaluationMore projects & source code | |
| DoomSat — READMEgithub.com | 2 | A DoomSat mission-stack experiment has Jev choose the spacecraft game pilot’s next intent; without a TypeSafe key, the README offers manual play. | Benchmarks & evaluationMore projects & source code | |
| dsh-jev-lensgithub.com | 0 | Jev quality bench for DeepSeek Harness: shadow-accounted dangerous-command judging, injection screening, canary A/B drills, optional gate mode. | Benchmarks & evaluationMore projects & source code | |
| dsh-jev-lens — READMEgithub.com | 0 | DSH experimental plugin measures Jev’s effect on agent task quality through accounting, screening and poisoning controls, with blocking enabled only in gate mode. | Benchmarks & evaluationMore projects & source code | |
| dsh-jev-verify — READMEgithub.com | 1 | The README presents Jev/System One as a first-class DeepSeek Harness plugin for typed decisions and verification. | Benchmarks & evaluationMore projects & source code | |
| EmreKaplaner/rag-jevgithub.com | 2 | Make room for useful evidence. Inspectable context selection for RAG, with Jev reranking and open benchmark studies. | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 2 | 4anti/jev-broadcast-lab — Operator lab for TypeSafe Jev. Chess Arena, closed-schema booths, Stockfish HUD for review only. · endpoint · JavaScript | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | AnshChoudhary/typesafe-ai-firewall — Shadow-mode validation harness for a pre-execution firewall on AI agent tool calls (TypeSafe/Jev). Real run, findings in report.md. · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 2 | CompleteTech-LLC-AI-Research/jev-311-heatmap — NYC 311 complaint heatmaps with TypeSafe JEV: reproducible pipeline, live research results, and interactive geographic visualizations. · endpoint · HTML | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 2 | OmarMujahid/jev-decision-bench — An independent benchmark of TypeSafe's Jev, a model that does not write text. You send it some content and a list of typed questions (yes/no, pick one option, rate on a scale) and · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 2 | PistachioAIHQ/jev-synergy-screening — Jev (TypeSafe System One) × ASReview SYNERGY abstract screening demo — Choice/Noul vs gold labels · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 3 | RINNECODER/jev-behavior-study — Independent Jev 1.13.0 behavior study: report, controlled prompt experiments, raw results, and offline verification. · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 3 | SamuelSacco/jev-exploration — Jev (TypeSafe) exploratory thread: claim audit, live demos, and runnable code · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 9 | Tech-Byte-Frontier/jevgate — File-scoped maintainability review with TypeSafe Jev · endpoint · Rust | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 4 | Thanh-Mathieu95/jev-model-tokengate — Every token passes the gate before the screen. · endpoint · JavaScript | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 3 | TokenTrim/jev-agent-failure-benchmark — Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset). · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 2 | TokenTrim/jev-routing-experiment — Benchmarking TypeSafe's Jev decision model as a cost-efficient LLM router on RouterArena · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 8 | aabolfazl/typesafe-local — Inspired by TypeSafe Ai, Ask a local LLM typed questions, get calibrated probabilities instead of text. Structured output without generation or parsing. MLX / Apple Silicon. · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | acoyfellow/nightglass — Owned, deterministic classifier for checking whether agent claims are supported by evidence, with optional Jev comparison through Cloudflare AI Gateway. · endpoint · JavaScript | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 7 | anisselbd/jev-phishing-bench — Jev (TypeSafe) vs Claude Haiku 4.5 on 2 000 phishing emails: accuracy, calibration, latency, cost. Reproducible benchmark. · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 7 | arunav25/jev-mcp — Connect JEV to MCP clients and compare its judgments against general-purpose LLMs using shared datasets and measurable accuracy. · endpoint · JavaScript | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | copyleftdev/jev-labs — Never confidently wrong: a TLA+-verified consensus kernel around TypeSafe's Jev, run through 1,680 chaos-tested pharmacy decisions with zero wrong verdicts. Film, code, and every captured call. · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 192 | fstandhartinger/jevbench — JevBench v1 - a benchmark for Jev-class typed decision models: smart, cheap, fast, reliable, open. · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 5 | gemanor/jev-code-review-benchmark — Comparing Jev, Gemini Flash, and Claude Fable on Python code review rules: cost, speed, accuracy, and consistency. Includes results, charts, and reproducible experiments. · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | integrate-your-mind/jev-nethack — Jev x NetHack: bounded runner, research code, and completed recording releases · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 6 | lbotinelly/jev-little-airways — A show-and-tell capability study for Jev, TypeSafe's System One decision model. · endpoint · HTML | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | markfive-proto/typesafe-vs-deepseek — TypeSafe (Jev) vs DeepSeek-flash: side-by-side speed/token/cost/accuracy comparison across invoice extraction, email classification, and reranking · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 4 | markjaquith/typesafe-ai-playground — A playground for experiments around Jev, TypeSafe's System One model. · endpoint · Rust | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 19 | mizchi/jev-playground — TypeSafe AI の System One モデル Jev を MoonBit から触るためのプレイグラウンド。 · endpoint · TypeScript | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | pavy23/jev_typesafeai_test — TypeSafe AI (jev) 시범 사용 프로젝트 · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | pokertools-arena/pokertools-arena.github.io — A browser-first AI poker benchmark. Seat Jev and OpenAI-compatible models at the same no-limit Texas Hold'em table, watch every card and decision as a spectator, and let the tournament run autonomously until one model wins. · endpoint · JavaScript | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | rogeriochaves/jev-experiments — Results page: https://claude.ai/code/artifact/f00ee126-9554-4e2f-b2e7-1fc86c066aa9 · endpoint · Go | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | ruffood/jev-reality-check — Jev (TypeSafe AI) 可复现实测:算术、计数、日期、零幻觉、把握度校准,中英对照 · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | simonmesmith/jev-bbq-experiment — Reproducible evaluation of TypeSafe Jev on all 58,492 BBQ questions: accuracy, stereotype bias, uncertainty, cost and latency. · endpoint · R | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 2 | sysadarsh/zerosweep — Autonomous System-One Triage Engine & Benchmark powered by TypeSafe AI (Jev). 75ms inference, $0 output tokens, and RLCD epistemic safety gates. · endpoint · TypeScript | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | tedliou/decision-model-playground — A local playground for comparing Laya and Jev decision models with article recommendations. · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 7 | thejorgg/omp-jev — TypeSafe Jev routing for Oh My Pi, with an opt-in checkpoint orchestrator and editable XDG configuration. Requires Bun = 1.3.14 and OMP = 18.2.3. · endpoint · TypeScript | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 6 | y0usaf/jev-lm — A word-level language model whose output layer is Jev: n-gram drafter, Noul chunk verification, bits-per-token eval · endpoint · TypeScript | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | zsavage8/padflow-jev-evals — Typed-decision benchmark from PadFlow (land development SaaS): schemas, anonymized labeled rows, and a runner for confidence-calibrated models like TypeSafe Jev. · endpoint · Python | Benchmarks & evaluationMore projects & source code | |
| endpointgithub.com | 1 | zzzzzec/jevsort — Jev-powered integer sorting experiment: serial selection versus parallel rank prediction. · endpoint · HTML | Benchmarks & evaluationMore projects & source code | |
| english-2-sqlgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| english-2-sql — READMEgithub.com | 0 | This learning scaffold asks Jev for semantic query-plan choices, compiles them to a SQL preview, and documents that no PostgreSQL connection exists yet. | Benchmarks & evaluationMore projects & source code | |
| Ensemblr Jev decision layer studygithub.com | — | Design proposal and spike for using Jev in a desktop orchestrator for Pi and Claude Code to pick agent roles, rate difficulty and flag duplicates; rejected after a corrected rerun that still did not support production use. | Benchmarks & evaluationMore projects & source code | |
| Ensemblr Jev decision layer study — repogithub.com | — | Design proposal and spike for using Jev in a desktop orchestrator for Pi and Claude Code to pick agent roles, rate difficulty and flag duplicates; rejected after a corrected rerun that still did not support production use. | Benchmarks & evaluationMore projects & source code | |
| Ensemblr Jev decision layer study — runbookgithub.com | — | Design proposal and spike for using Jev in a desktop orchestrator for Pi and Claude Code to pick agent roles, rate difficulty and flag duplicates; rejected after a corrected rerun that still did not support production use. | Benchmarks & evaluationMore projects & source code | |
| estudo-system-onegithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| estudo-system-one — READMEgithub.com | 0 | Portuguese study materials cover Jev’s typed decision primitives and include Kev practice and calibration exercises. | Benchmarks & evaluationMore projects & source code | |
| Evaluating jevgithub.com | — | Pre-registered adversarial evaluation of jev-1.13.0 with nine experiments and 28 predictions fixed before any data, run as 123,805 requests for $12.69, plus a 13-rule prompting guide drawn from the results. | Benchmarks & evaluationMore projects & source code | |
| fab-evidence-gategithub.com | 0 | Evidence-aware semiconductor alert triage research demo with TypeSafe Jev, policy guards, and reproducible evaluation | Benchmarks & evaluationMore projects & source code | |
| fab-evidence-gate — READMEgithub.com | 0 | Synthetic factory-alert demo uses Jev to propose review routes, urgency and needed evidence after deterministic high-risk screening; humans retain final control. | Benchmarks & evaluationMore projects & source code | |
| fabricioism/jev-expirementsgithub.com | 1 | A repository for jev-expirements | Benchmarks & evaluationMore projects & source code | |
| FinancialPredictionJevgithub.com | 0 | Single Python script that asks a Jev Noul whether SPY closes up the next day from yfinance data, then scores accuracy, Brier score, ROC-AUC and calibration against a baseline. | Benchmarks & evaluationMore projects & source code | |
| FinancialPredictionJev — untitled15.pygithub.com | 0 | A Python finance experiment sends market-state features to TypeSafe System One and asks a Noul question about the probability of a positive next-day close-to-close return; no prediction results were verified here. | Benchmarks & evaluationMore projects & source code | |
| FinancialPredictionJev — untitled15.pygithub.com | 0 | A Python finance experiment sends market-state features to TypeSafe System One and asks a Noul question about the probability of a positive next-day close-to-close return; no prediction results were verified here. | Benchmarks & evaluationMore projects & source code | |
| finetuningsingh/jev-chatbotgithub.com | 1 | Experiment: using TypeSafe Jev as a chatbot by choosing replies one letter or word at a time | Benchmarks & evaluationMore projects & source code | |
| foreman-jev-evaluationgithub.com | 0 | JEV evaluation (FM-JEV-01): replay-only, advisory-only evaluation of Foreman-style supervision with TypeSafe Jev. Measures how much safety comes from the model versus a deterministic evidence gate. No worker authority. | Benchmarks & evaluationMore projects & source code | |
| foreman-jev-evaluation — READMEgithub.com | 0 | This replay-only, advisory-only study evaluates how TypeSafe Jev behaves when asked to supervise a coding-agent workflow. | Benchmarks & evaluationMore projects & source code | |
| Formagithub.com | 0 | Design harness where Jev picks a page, screen, form, or diagram design in about a second from your description and an LLM, Luna, only fills in the words, with a switch to see Jev alone. | Benchmarks & evaluationMore projects & source code | |
| forma-system1-experiment — READMEgithub.com | 0 | Forma is an experiment where Jev chooses page structure and code assembles the design; the README describes it as an experiment, not a product. | Benchmarks & evaluationMore projects & source code | |
| fraud-jevgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| fraud-jev — READMEgithub.com | 0 | This Next.js demo compares TypeSafe Jev with a Claude path on fictional payment cases, not real fraud outcomes. | Benchmarks & evaluationMore projects & source code | |
| ghost-usergithub.com | 1 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| ghost-user — READMEgithub.com | 1 | The browser-session helper gives Jev a compact page state and asks for one typed next action. | Benchmarks & evaluationMore projects & source code | |
| GoodWatch Jev fingerprint experimentgithub.com | — | Case study from the GoodWatch movie-discovery app testing Jev for 74 per-title trait scores: 10 to 25 times faster but a 2.65-point mean gap to reviewed Qwen scores, so the team decided not to adopt it. | Benchmarks & evaluationMore projects & source code | |
| GoodWatch Jev fingerprint experiment — adrgithub.com | — | Case study from the GoodWatch movie-discovery app testing Jev for 74 per-title trait scores: 10 to 25 times faster but a 2.65-point mean gap to reviewed Qwen scores, so the team decided not to adopt it. | Benchmarks & evaluationMore projects & source code | |
| GoodWatch Jev fingerprint experiment — repogithub.com | — | Case study from the GoodWatch movie-discovery app testing Jev for 74 per-title trait scores: 10 to 25 times faster but a 2.65-point mean gap to reviewed Qwen scores, so the team decided not to adopt it. | Benchmarks & evaluationMore projects & source code | |
| got-jev — READMEgithub.com | 1 | In this role-play demo, another model writes scenes and TypeSafe Jev labels each scene with typed answers; the README explicitly says Jev does not write the story. | Benchmarks & evaluationMore projects & source code | |
| GPT vs JEVgithub.com | 1 | Side-by-side learning demo that sends the same input to GPT for free text and to Jev for structured Noul probabilities, with saved example runs and limited live comparisons. | Benchmarks & evaluationMore projects & source code | |
| gpt-vs-jev — READMEgithub.com | 1 | The comparison demo contrasts GPT text generation with Jev's structured Noul yes-probability output. | Benchmarks & evaluationMore projects & source code | |
| ground-zero — READMEgithub.com | 1 | Ground Zero describes TypeSafe Jev as a decision capability for evaluating model responses against supplied criteria. | Benchmarks & evaluationMore projects & source code | |
| HAKARI-Bench Jev rerankergithub.com | — | Jev integration in HAKARI-Bench, a lightweight IR benchmark over 35+ benchmark groups, that ranks documents by Noul relevance probabilities in listwise or pointwise mode with jev-1.13.0 pinned. | Benchmarks & evaluationMore projects & source code | |
| HAKARI-Bench Jev reranker — repogithub.com | — | Jev integration in HAKARI-Bench, a lightweight IR benchmark over 35+ benchmark groups, that ranks documents by Noul relevance probabilities in listwise or pointwise mode with jev-1.13.0 pinned. | Benchmarks & evaluationMore projects & source code | |
| hellow-jevgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| hellow-jev — READMEgithub.com | 0 | Classification benchmark compares hosted Jev/Claude and local Laya/Qwen approaches, including e-commerce logs and standard text-classification datasets. | Benchmarks & evaluationMore projects & source code | |
| here-we-go-jev — READMEgithub.com | 1 | A Jev playground/test bench for state and typed questions with mock and System One backends. Mock runs without a key and must not be treated as live model evidence. | Benchmarks & evaluationMore projects & source code | |
| Hermes Agent compaction scorecardgithub.com | — | Tests a Jev-based context-compaction plugin against Hermes' own compressor and recommends against it: far cheaper and faster per compaction, but it kept twice the context and ranked tool results no better than recency. | Benchmarks & evaluationMore projects & source code | |
| i-shoot-rockgithub.com | 0 | jev-based rock paper scissors, to test jev pre-trained model. | Benchmarks & evaluationMore projects & source code | |
| i-shoot-rock — READMEgithub.com | 0 | Rock-paper-scissors game gives Jev only previous rounds as state to predict the next move; the request starts before the player's current move and code resolves the game. | Benchmarks & evaluationMore projects & source code | |
| integrate-your-mind/jev-nethackgithub.com | 1 | Jev x NetHack: bounded runner, research code, and completed recording releases | Benchmarks & evaluationMore projects & source code | |
| Is Jev cheaper and better?github.com | — | Reproducible invoice-classification test on 50 documents built so surface cues mislead: Jev scored 50/50 at $0.025 per 1,000 decisions, tied with Claude Haiku 4.5, while two open models scored 48/50. | Benchmarks & evaluationMore projects & source code | |
| Is Jev cheaper and better? — repogithub.com | — | Reproducible invoice-classification test on 50 documents built so surface cues mislead: Jev scored 50/50 at $0.025 per 1,000 decisions, tied with Claude Haiku 4.5, while two open models scored 48/50. | Benchmarks & evaluationMore projects & source code | |
| jaggedgithub.com | 0 | A pre-registered test of which parts of TypeSafe's advice for its Jev model change the answer. | Benchmarks & evaluationMore projects & source code | |
| jagged — READMEgithub.com | 0 | This preregistered research varies TypeSafe’s recommendations on the Wikipedia Articles for Deletion task; reported results remain the authors’ experiment, not a reproduced finding. | Benchmarks & evaluationMore projects & source code | |
| Janusgithub.com | 2 | Router that sends each decision to Jev or a larger fallback model based on confidence, after measuring the threshold on your labeled dataset or decision log instead of shipping a default. | Benchmarks & evaluationMore projects & source code | |
| Jebadiahgithub.com | — | Open replica: Apache-2.0 decision models (27B, 9B, 4B on Qwen bases; bf16, GGUF and MLX) that answer Choice, Noul and Score questions with a probability for every option from one forward pass, and run anywhere: a standalone server with Jev's /v1/systemone wire and a playground, a llama.cpp script for the GGUF builds, or AINode (open-source local AI platform). | Benchmarks & evaluationMore projects & source code | |
| JEVgithub.com | 2 | 调研报告:TypeSafe System One 决策模型与开源对标(含原始核实记录) | Benchmarks & evaluationMore projects & source code | |
| jevgithub.com | 0 | Independent research notes toward an open Jev-like decision model: public facts, API contract, training and eval plan. | Benchmarks & evaluationMore projects & source code | |
| Jev 1.13 as a Reward Modelgithub.com | 3 | Reproducible evaluation of Jev 1.13 as a reward model, LLM judge and process verifier across eight benchmark tracks and 40,940 examples, with an interactive report of 54 SOTA comparisons. | Benchmarks & evaluationMore projects & source code | |
| Jev 2048github.com | 5 | Instrumented 2048 web lab where every move is a Jev Choice over four directions with no heuristic fallback, showing probabilities, confidence, latency and cost live. | Benchmarks & evaluationMore projects & source code | |
| Jev Atari Labgithub.com | 0 | Research lab that plays Arcade Learning Environment Atari games with Jev structured decisions and value questions while a teacher edits the question program, publishing inputs, actions, rewards, and videos. | Benchmarks & evaluationMore projects & source code | |
| Jev Capability Atlasgithub.com | 26 | Bilingual, mostly Traditional Chinese evidence map of where Jev's calibrated-decision claim holds and where it breaks, built from real API-call receipts, test suites, and guides for agents. | Benchmarks & evaluationMore projects & source code | |
| Jev Column Racegithub.com | 23 | Live race labeling 1,000 app reviews in four columns: Jev finished in 4.6 seconds for $0.023 versus 18.8 seconds and $0.158 for Gemini 3.8 Flash, with similar star-rating agreement. | Benchmarks & evaluationMore projects & source code | |
| Jev in Koreangithub.com | 6 | Frozen, reproducible 100-question sample check of Jev on Korean text: reading comprehension scored 96 in Korean versus 97 in English, while fine-grained meaning judgments scored 76 versus 80. | Benchmarks & evaluationMore projects & source code | |
| Jev lab (TypeScript CLI)github.com | 0 | Small Node.js TypeScript CLI that tries Jev on synthetic support text: one urgency question, three questions in a single request, and a latency comparison between them. | Benchmarks & evaluationMore projects & source code | |
| Jev pick-and-place studygithub.com | 1 | Reproducible MuJoCo pilot comparing Jev 1.13.0, Claude Haiku 4.5 and reactive rules as pick-and-place controllers; Jev succeeded 10/10 in both settings at 146 ms median per decision, matching the rules exactly. | Benchmarks & evaluationMore projects & source code | |
| Jev Playgroundgithub.com | 1 | Benchmark playground that pits Jev against Luna, Haiku and Gemini at choosing moves in explicit-state games, where game code owns the rules and legal actions and each model only picks among them. | Benchmarks & evaluationMore projects & source code | |
| Jev Ponggithub.com | — | Browser Pong where the ball moves one step per model decision, pitting Jev against LLMs through Vercel AI Gateway, with every player and agent connected over an Ably channel. | Benchmarks & evaluationMore projects & source code | |
| Jev research deckgithub.com | 0 | Ongoing Japanese research on Jev and System One models maintained as Markdown slides, a presentation script and an article, backed by a ledger of sources, third-party verification and caveats. | Benchmarks & evaluationMore projects & source code | |
| Jev research paper packagegithub.com | 1 | Reproducible research package on improving Jev decisions: entity-matching macro-F1 rose from 0.9605 to 0.9859 on 413 DBLP-ACM pairs, while SciFact relation verification showed no resolved gain. | Benchmarks & evaluationMore projects & source code | |
| Jev synthetic surveygithub.com | 2 | Study running Jev and GPT-4.1 as the same 300 synthetic respondents over 24,596 Twin-2K-500 cells, finding that asking yes/no items as a Noul mattered more than the model gap, at a thirty-fourth of the cost. | Benchmarks & evaluationMore projects & source code | |
| Jev × Cohen ADHD Abstract Triagegithub.com | 2 | Systematic-review screening demo that has Jev include or exclude MEDLINE abstracts for an ADHD drug review; on all 851 Cohen et al. 2006 abstracts it reached 83.3% recall and 68.6% precision, tuned on that same set. | Benchmarks & evaluationMore projects & source code | |
| jev — READMEgithub.com | 0 | These open notes collect information on TypeSafe Jev and outline a plan for an equivalent open decision engine; they do not present that engine as already built. | Benchmarks & evaluationMore projects & source code | |
| JEV — READMEgithub.com | 2 | This repository is a research report about Jev/System One and alternative decision-model projects, not an implementation. | Benchmarks & evaluationMore projects & source code | |
| Jev-2048github.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-2048 — READMEgithub.com | 5 | A web lab delegates every 2048 move to Jev Choice and surfaces probabilities, confidence and latency; the README explicitly says it has no heuristic fallback. | Benchmarks & evaluationMore projects & source code | |
| Jev-2048 — readmegithub.com | 0 | 2048 experiment compares Jev play with different board, history and feature-state inputs against deterministic baselines, with a dashboard for automatic or manual play. | Benchmarks & evaluationMore projects & source code | |
| jev-access-daygithub.com | 0 | A learning scaffold for TypeSafe AI's System One models: eval harness plus a measured, plain-language comparison of the Jev decision model vs an LLM stand-in on 24 real operational decisions. All numbers reproducible from committed run files. | Benchmarks & evaluationMore projects & source code | |
| jev-access-day — READMEgithub.com | 0 | typesafe-lab is a learning/evaluation scaffold with a recorded live Jev comparison; its author explicitly says it is not a benchmark. | Benchmarks & evaluationMore projects & source code | |
| jev-acentogithub.com | 0 | Pre-registered audit of Jev on Spanish over 3,200 paired human-labeled items: a Spanish state cost accuracy on every dataset, up to 6.4 pp on XNLI, while Spanish instructions changed nothing; includes a CLI to rerun it. | Benchmarks & evaluationMore projects & source code | |
| jev-acento — READMEgithub.com | 0 | jev-acento is an independent reproducible audit of Jev on Spanish data, with a CLI for repeating the comparison on labeled data. | Benchmarks & evaluationMore projects & source code | |
| jev-agent-failure-benchmarkgithub.com | 3 | Benchmark of Jev on the 6,257 text traces of Who&When Pro, attributing multi-agent failures to an agent, step and error type; Jev beat GPT-5.4 on every axis for ~$1.28 total. | Benchmarks & evaluationMore projects & source code | |
| jev-agent-failure-benchmark — READMEgithub.com | 3 | A repository benchmarks Jev against an LLM on the text subset of an agent-failure-attribution dataset; its reported scores are the project’s experiment, not independent validation. | Benchmarks & evaluationMore projects & source code | |
| jev-agent-safetygithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-agent-safety — READMEgithub.com | 0 | An exploratory study applies Jev questions to 108 published AI-incident excerpts; the author says its observations are not validated accuracy or evidence of deployment safety. | Benchmarks & evaluationMore projects & source code | |
| jev-alpha-benchgithub.com | 0 | Two studies on whether Jev predicts stock returns from news headlines or price candles: same-day rank IC of +0.24 falls to -0.008 by the next close, and a long-short book loses 18 bps a trade after costs. | Benchmarks & evaluationMore projects & source code | |
| jev-alpha-bench — READMEgithub.com | 0 | This benchmark project asks whether TypeSafe Jev predicts stock returns from financial news; it is a research test, not a trading product. | Benchmarks & evaluationMore projects & source code | |
| jev-anotacao-sentencas — READMEgithub.com | 0 | A research experiment compares Jev with Gemini and GPT on structured annotation of 120 civil judgments from São Paulo state court; reported results were not re-run here. | Benchmarks & evaluationMore projects & source code | |
| Jev-api-experimentsgithub.com | 0 | Empirical experiments and API research for TypeSafe's Jev System One model trained using RLCD | Benchmarks & evaluationMore projects & source code | |
| Jev-api-experiments — READMEgithub.com | 0 | This research repository runs controlled Jev API experiments, preserves raw evidence, and updates hypotheses only after measurement. | Benchmarks & evaluationMore projects & source code | |
| jev-applicationgithub.com | 0 | SInce Jev model is popular, let's try some interesting. | Benchmarks & evaluationMore projects & source code | |
| jev-application — READMEgithub.com | 0 | Reworks five classic ML classification tasks with Jev, comparing it with traditional models and exploring few-shot examples and stacking. | Benchmarks & evaluationMore projects & source code | |
| jev-arcadegithub.com | 0 | Can a System One model play arcade games? TypeSafe's Jev plays Tetris, Snake and 2048 — benchmarked against random and heuristic baselines. | Benchmarks & evaluationMore projects & source code | |
| jev-arcade — READMEgithub.com | 0 | This arcade-game experiment explores Jev action choices across games such as Tetris and Sokoban; any reported results are the authors’ experiments, not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-as-a-judge — READMEgithub.com | 97 | This research project compares Jev with generative-model judges on fixed agent runs using accuracy, score reliability, cost and latency measures; results are not rerun here. | Benchmarks & evaluationMore projects & source code | |
| jev-as-llmgithub.com | 0 | TypeSafe's Jev, a judgment model that was never trained to write, made to chat one word at a time. Browser-only, bring your own OpenRouter key. | Benchmarks & evaluationMore projects & source code | |
| jev-as-llm — READMEgithub.com | 0 | Experimental chat assembles text word by word through repeated Jev multiple-choice calls, showing alternative probabilities and replaying recorded runs without a key. | Benchmarks & evaluationMore projects & source code | |
| Jev-as-Policy — READMEgithub.com | 47 | A MuJoCo manipulation demo uses Jev for sequential intent and motor-choice judgments; geometry, inverse kinematics and physical control remain in the local harness. | Benchmarks & evaluationMore projects & source code | |
| jev-atari-lab — READMEgithub.com | 0 | A research lab studies Jev question policies for Atari agents and publishes experiment evidence; its README says no repeatable optimization benefit has been established. | Benchmarks & evaluationMore projects & source code | |
| jev-banc-francais — READMEgithub.com | 0 | Jev Banc Français is a provider-neutral harness for evaluating closed-set decisions on French-language cases. | Benchmarks & evaluationMore projects & source code | |
| jev-baselines-eval — READMEgithub.com | 5 | Independent pre-registered studies of TypeSafe Jev intent classification and confidence-gated cascades. Listed for evaluation evidence and baseline methodology, not validation of every published result. | Benchmarks & evaluationMore projects & source code | |
| jev-battleshipgithub.com | 1 | Battleship against Jev, a model that answers in probabilities instead of text. Web game plus a CLI arena that plays it against general-purpose LLMs on identical fleets. | Benchmarks & evaluationMore projects & source code | |
| jev-battleship — READMEgithub.com | 1 | A Battleship web game lets a person play against Jev and includes a CLI arena that compares Jev with a general-purpose LLM on identical fleets. | Benchmarks & evaluationMore projects & source code | |
| jev-behavior-studygithub.com | 3 | Independent field guide to Jev 1.13.0 behavior from 11,621 text-study requests, 3 Snake studies, and a 3D City lab, showing where framing changes answers and harder tasks fail. | Benchmarks & evaluationMore projects & source code | |
| jev-behavior-study — READMEgithub.com | 3 | This field guide reports task-specific observations from Jev 1.13.0 studies and explicitly says they are not an official benchmark or overall model score. | Benchmarks & evaluationMore projects & source code | |
| jev-benchgithub.com | 0 | Measure accuracy and calibration of Jev (TypeSafe AI's decision model) on public datasets: 12 business-like tasks, 7 experiments, one Python file. | Benchmarks & evaluationMore projects & source code | |
| jev-benchgithub.com | 1 | Does the cited source actually say it? A 42-claim benchmark: Jev (TypeSafe System One) against GPT-5.4, Claude Sonnet 5 and Gemini 3.1 Pro. | Benchmarks & evaluationMore projects & source code | |
| jev-benchgithub.com | 0 | TypeSafe / Jev community project: thomasschafer/jev-bench. | Benchmarks & evaluationMore projects & source code | |
| jev-bench — READMEgithub.com | 1 | This project benchmarks the narrow task of judging whether a cited passage supports a claim sentence; it is a Jev-specific comparison study. | Benchmarks & evaluationMore projects & source code | |
| jev-bench — READMEgithub.com | 0 | jev-bench is an independent accuracy/calibration benchmark on public datasets; its README emphasizes that confidence behavior differs across tasks. | Benchmarks & evaluationMore projects & source code | |
| jev-benchmarkgithub.com | 0 | 面向语义决策模型的中文言下之意评测集:100 道伴侣对话 Choice 题,支持本地 ONNX、Jev 官方 API 与内网模型对比 | Benchmarks & evaluationMore projects & source code | |
| jev-benchmarkgithub.com | 1 | Reproducible benchmark of Jev classifying agent tool calls as readonly, destructive, privileged or exfiltration, measuring accuracy, latency and whether its confidence is worth routing on, with raw results committed. | Benchmarks & evaluationMore projects & source code | |
| jev-benchmarkgithub.com | 6 | Two Jev benchmarks: chess, where it plays at roughly 950 Elo only with code-supplied tactical facts, and deciding which game NPC a speech-to-text player is addressing, at F1 0.96 with precision 1.0. | Benchmarks & evaluationMore projects & source code | |
| jev-benchmark — READMEgithub.com | 0 | Chinese intent benchmark compares hosted Jev with local Jev-like models on 100 multi-turn relationship scenarios using four-option Choice questions. | Benchmarks & evaluationMore projects & source code | |
| jev-benchmark — READMEgithub.com | 6 | The playground describes hands-on Jev benchmarks on chess and NPC addressee detection, with application code owning the workflow. | Benchmarks & evaluationMore projects & source code | |
| jev-benchmarks — READMEgithub.com | 21 | The benchmark README describes a comparison of TypeSafe Jev with another classifier on zero-shot single-label decisions and per-label probabilities. | Benchmarks & evaluationMore projects & source code | |
| jev-bias-benchgithub.com | 1 | A benchmark for Jev's biases, built from people who differ in one attribute at a time. | Benchmarks & evaluationMore projects & source code | |
| jev-bias-bench — READMEgithub.com | 1 | This independent study changes one person attribute at a time, sends the same case to Jev through the PHP SDK, and compares the judgments. | Benchmarks & evaluationMore projects & source code | |
| jev-btzscgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-btzsc — READMEgithub.com | 0 | jev-btzsc is a small harness for strict zero-shot text classification using TypeSafe Jev, with one request per source example; reported scores are not repeated. | Benchmarks & evaluationMore projects & source code | |
| jev-capability-atlas — READMEgithub.com | 26 | This independent community atlas documents tasks that may or may not fit Jev, rather than providing a runtime integration. | Benchmarks & evaluationMore projects & source code | |
| jev-certify — READMEgithub.com | 1 | A research tool applies conformal risk-control methods to turn Jev choices into routing thresholds; its reported experiment is based on logged Jev decisions. | Benchmarks & evaluationMore projects & source code | |
| jev-chatgithub.com | 4 | Research decoder that builds a chatbot from Jev Choices over a hierarchical codebook of phrases and words, with a paper comparing stepwise decoding against selecting a complete reply. | Benchmarks & evaluationMore projects & source code | |
| jev-classification-promptinggithub.com | 0 | Reproducible experiments on prompting criteria in TypeSafe's Jev — how much of a classifier's behaviour is yours to define | Benchmarks & evaluationMore projects & source code | |
| jev-classification-prompting — READMEgithub.com | 0 | An exploratory study tests how criteria prompts affect TypeSafe Jev’s typed classification behavior; the README labels the work ongoing and its results model-version-specific. | Benchmarks & evaluationMore projects & source code | |
| jev-codebookgithub.com | 0 | Qualitative coding at scale with Jev: apply a codebook to open-ended text, review uncertain items, measure agreement against your human coders. | Benchmarks & evaluationMore projects & source code | |
| jev-codebook — READMEgithub.com | 0 | jev-codebook applies user-defined qualitative labels to survey, review, or interview text using Jev. | Benchmarks & evaluationMore projects & source code | |
| jev-column-race — READMEgithub.com | 23 | A live comparison project races Jev against Gemini 3.8 Flash on a stated set of 1,000 app reviews; the README report is not independently reproduced. | Benchmarks & evaluationMore projects & source code | |
| jev-computergithub.com | 0 | An 8-bit computer built from one yes/no question asked to Jev (TypeSafe) — 24,511 NAND gates from a single API call | Benchmarks & evaluationMore projects & source code | |
| jev-computer — READMEgithub.com | 0 | jev-computer is an educational logic-computer demonstration mapping Jev answers to a NAND gate; it does not establish Jev as a general-purpose computer or hardware controller. | Benchmarks & evaluationMore projects & source code | |
| jev-consumer-researchgithub.com | 0 | Consumer research explorations using TypeSafe's Jev. | Benchmarks & evaluationMore projects & source code | |
| jev-contract-graphgithub.com | 0 | Conditional payoff proofs for prediction-market contracts with optional Jev semantic review. | Benchmarks & evaluationMore projects & source code | |
| jev-contract-graph — READMEgithub.com | 0 | Read-only binary-contract research prototype analyzes implications and conditional payoff proofs; optional Jev judgments remain separate from data validation and financial execution. | Benchmarks & evaluationMore projects & source code | |
| jev-crowdsimgithub.com | 0 | Run a declared factorial audience grid through typed Jev reactions and expose disagreement. | Benchmarks & evaluationMore projects & source code | |
| jev-crowdsim — READMEgithub.com | 0 | jev-crowdsim evaluates a message across an explicit grid of declared traits; Jev returns finite reactions while code computes main effects and Wilson intervals. Synthetic fixtures are not measured people or Jev performance. | Benchmarks & evaluationMore projects & source code | |
| jev-cyrillic-auditgithub.com | 0 | Does TypeSafe's Jev keep its accuracy and calibration on Russian? Independent RU vs EN audit (ECE, reliability diagrams, paired bootstrap) on parallel human-labelled data. | Benchmarks & evaluationMore projects & source code | |
| jev-cyrillic-audit — READMEgithub.com | 0 | The README reports a preregistered, paired Jev evaluation on Russian versus English XNLI items and says Russian accuracy and calibration were lower; results were not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-decision-bench — READMEgithub.com | 2 | An evaluation harness for conventional language models on TypeSafe Jev-style bounded-decision tasks, not a general text-generation leaderboard or proof of hosted Jev performance. | Benchmarks & evaluationMore projects & source code | |
| jev-deferred-crispification — READMEgithub.com | 0 | This repository presents a research preprint about calibration across multi-step Jev-style decision pipelines and proposes “Deferred Crispification”; it is analysis, not model weights. | Benchmarks & evaluationMore projects & source code | |
| jev-demogithub.com | 1 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-demogithub.com | 0 | Historical paper-trading simulator for evaluating TypeSafe AI JEV decisions | Benchmarks & evaluationMore projects & source code | |
| jev-demogithub.com | 0 | Decision Arena: TypeSafe's Jev vs open-source Laya playing highway-env, Snake and Blackjack with zero training, plus benchmarks and a Claude Code watchdog | Benchmarks & evaluationMore projects & source code | |
| jev-demo — READMEgithub.com | 1 | A Rust demo layers Jev judgments over deterministic trading and optimization examples; its README says the data is synthetic and no real market is connected. | Benchmarks & evaluationMore projects & source code | |
| jev-demo — READMEgithub.com | 0 | An educational paper-trading simulator compares Jev-directed stock decisions with baselines on completed historical U.S. trading days; it is not a live-return claim. | Benchmarks & evaluationMore projects & source code | |
| jev-demo — READMEgithub.com | 0 | Next.js evaluation lab provides Noul, custom Choice and ordinal Score tabs over a shared background, storing drafts, language settings and the API key in the browser. | Benchmarks & evaluationMore projects & source code | |
| jev-demo — READMEgithub.com | 0 | This repository includes Jev-vs-Laya arenas for highway-env, Snake, and Blackjack with requests and answers displayed; game and benchmark results are not independently evaluated here. | Benchmarks & evaluationMore projects & source code | |
| jev-demo — READMEgithub.com | 0 | This independent demo compares Jev and an OpenAI model on support-ticket triage for a fictional invoicing service; it does not establish real-world quality. | Benchmarks & evaluationMore projects & source code | |
| jev-demosgithub.com | 0 | Dart maze experiments testing Jev's spatial lookahead: asked for up to 100 future moves per request it solved zero mazes, but with adjacent-tile hints and one next-move question it solved 6/10 5x5 mazes. | Benchmarks & evaluationMore projects & source code | |
| jev-demos — READMEgithub.com | 0 | This repository collects demos that use TypeSafe AI System One models. | Benchmarks & evaluationMore projects & source code | |
| jev-dev — READMEgithub.com | 0 | This viewer sends the same user utterance to TypeSafe Jev and an LLM for parallel judgments, then displays their differing affect states. | Benchmarks & evaluationMore projects & source code | |
| jev-divination-labgithub.com | 0 | Researching shared-axis comparison of independent divination readings with Jev / TypeSafe AI. | Benchmarks & evaluationMore projects & source code | |
| jev-divination-lab — READMEgithub.com | 0 | Research starter preserves a Jev connectivity example for planned comparisons of divination outputs; divination logic and daily workflows are not implemented yet. | Benchmarks & evaluationMore projects & source code | |
| jev-doom — READMEgithub.com | 2 | A Freedoom demo visualizes Jev's structured game decisions in a local dashboard and documents TypeSafe-direct or Vercel Gateway access. | Benchmarks & evaluationMore projects & source code | |
| jev-drivegithub.com | 1 | Jev + autonomous driving: structured decisions, multimodal baselines, recovery research, and measured API diagnostics. | Benchmarks & evaluationMore projects & source code | |
| jev-drive — READMEgithub.com | 1 | JevDrive is a research prototype for Jev recovery/replanning choices, while a local controller and feasibility checks own vehicle execution; closed-loop improvement is not demonstrated. | Benchmarks & evaluationMore projects & source code | |
| jev-drone — READMEgithub.com | 229 | A MuJoCo quadrotor simulation uses Jev to interpret camera-state situations at a documented cadence; ordinary code handles time-critical control. | Benchmarks & evaluationMore projects & source code | |
| jev-dspy-labgithub.com | 10 | Companion lab for DSPy pipelines that records and replays TypeSafe calls to measure Jev calibration, selective risk, confidence-gated abstention, latency, tokens, and modeled cost offline. | Benchmarks & evaluationMore projects & source code | |
| jev-dspy-lab — READMEgithub.com | 10 | Jev DSPy Lab is a companion measurement project for calibrating and confidence-gating Jev decisions in DSPy workflows, not a DSPy fork. | Benchmarks & evaluationMore projects & source code | |
| jev-enterprise-decision-fabricgithub.com | 0 | Architecture for running many semantic decisions through one validated path, with a labelled 111-case benchmark comparing TypeSafe Jev against a Claude baseline, and a dashboard for inspecting any single decision. Experimental, not production. | Benchmarks & evaluationMore projects & source code | |
| jev-enterprise-decision-fabric — READMEgithub.com | 0 | An experimental architecture describes using TypeSafe Jev for enterprise decision flows; the README explicitly says it is not production software. | Benchmarks & evaluationMore projects & source code | |
| jev-evalgithub.com | 0 | 在自己的数据上评测 TypeSafe Jev 的准确率、概率校准和可用阈值 | Benchmarks & evaluationMore projects & source code | |
| jev-eval — READMEgithub.com | 1 | A benchmark harness compares TypeSafe Jev with OpenRouter models or local checkpoints on labeled classification data. | Benchmarks & evaluationMore projects & source code | |
| jev-eval — READMEgithub.com | 0 | A Python tool evaluates TypeSafe Jev on user data for accuracy, probability calibration, and threshold selection. | Benchmarks & evaluationMore projects & source code | |
| jev-eval-agentgithub.com | 107 | Experiment with a personal-assistant agent and 100 mocked tools that counts the steps needed when the LLM picks tools itself versus when Jev picks the tool and the LLM only fills arguments. | Benchmarks & evaluationMore projects & source code | |
| jev-experimentsgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-experimentsgithub.com | 0 | Small experiments with Jev by TypeSafe | Benchmarks & evaluationMore projects & source code | |
| jev-experimentsgithub.com | 1 | experiments with system one model jev | Benchmarks & evaluationMore projects & source code | |
| jev-experimentsgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-experiments — mapreduce.mjsgithub.com | 0 | The repository’s map/reduce experiment replaces programmer-written reducers with Jev judgments and batches per-group questions into one System One request; this is a documented experiment, not a performance claim. | Benchmarks & evaluationMore projects & source code | |
| jev-experiments — READMEgithub.com | 396 | jev-experiments is a collection of TypeSafe/Jev latency-focused demos, with each demo in its own application directory. | Benchmarks & evaluationMore projects & source code | |
| jev-experiments — READMEgithub.com | 0 | A repository of small Jev experiments, including flight search, expense tagging, and news filtering; it is a demo collection, not one integrated production application. | Benchmarks & evaluationMore projects & source code | |
| jev-experiments — READMEgithub.com | 1 | Experiments comparing TypeSafe Jev with a reasoning model using identical machine-checkable tasks and metrics. Listed for evaluation artifacts, not endorsement of reported outcomes. | Benchmarks & evaluationMore projects & source code | |
| jev-experiments — READMEgithub.com | 0 | A TypeScript experiment suite for typed System One decisions lists TypeSafe-hosted Jev as the default backend for its use cases. | Benchmarks & evaluationMore projects & source code | |
| jev-exploration — READMEgithub.com | 3 | An evidence ledger tracks public claims about TypeSafe Jev, recording the status and source for each claim rather than providing an integration. | Benchmarks & evaluationMore projects & source code | |
| jev-fanout-benchgithub.com | 0 | Measures what a Jev request is billed and how its answers hold up under batching, translation, and rewording, from 3,455 billed requests with public raw logs. | Benchmarks & evaluationMore projects & source code | |
| jev-fanout-bench — READMEgithub.com | 0 | Benchmark measures Jev question fan-out across state sizes and question counts, comparing batched versus separate requests and testing the billing model. | Benchmarks & evaluationMore projects & source code | |
| jev-field-testsgithub.com | 0 | Twelve field tests for TypeSafe's Jev model: calibration, guardrails, résumé screening, interview rubrics and its failure modes. Bring your own API key. | Benchmarks & evaluationMore projects & source code | |
| jev-field-tests — READMEgithub.com | 0 | A research repository collects twelve exploratory Jev test scenarios, diagnostics, and saved evidence; the review does not treat its reported results as independently verified. | Benchmarks & evaluationMore projects & source code | |
| jev-field-trialgithub.com | 0 | A pre-registered field trial of Jev (TypeSafe's judgment model) on a second brain and Claude Code history: 20 tests, bars written first, failures included, and the tools to repeat it. | Benchmarks & evaluationMore projects & source code | |
| jev-field-trial — READMEgithub.com | 0 | The README documents a Jev field trial on the author’s personal second-brain system. | Benchmarks & evaluationMore projects & source code | |
| jev-finance-benchmarkgithub.com | 0 | Typesafe.ai model jev finance benchmark. | Benchmarks & evaluationMore projects & source code | |
| jev-freeformgithub.com | 3 | Observable chat experiment that makes Jev generate text one character at a time by choosing among 98 options (printable ASCII, newline, tab and end of sequence) for each next reply prefix. | Benchmarks & evaluationMore projects & source code | |
| jev-freeform — READMEgithub.com | 3 | jev-freeform explores whether successive character choices to TypeSafe Jev can produce freeform text; the README frames this as an experiment. | Benchmarks & evaluationMore projects & source code | |
| jev-graph-walkgithub.com | 0 | Context-carrying graph retrieval with Jev: branching walks, reproducible ablations, and a visual replay. | Benchmarks & evaluationMore projects & source code | |
| jev-graph-walk — READMEgithub.com | 0 | jev-graph-walk is an inspectable prototype for cascading graph retrieval with Jev; its browser viewer replays recorded calls on fictional graphs rather than invoking a model. | Benchmarks & evaluationMore projects & source code | |
| jev-graphrag — READMEgithub.com | 1 | A set of Neo4j GraphRAG demonstrations explores using Jev for judgment steps such as entity resolution and deduplication. | Benchmarks & evaluationMore projects & source code | |
| jev-guardbenchgithub.com | 0 | Can a System One model (TypeSafe Jev, open-source Kev) replace an LLM-as-judge in agent guardrail callbacks? | Benchmarks & evaluationMore projects & source code | |
| jev-guardbench — READMEgithub.com | 0 | jev-guardbench studies System One as an agent-guardrail judge, but README says self-hosted Kev-9B is primary and hosted Jev is a secondary arm. | Benchmarks & evaluationMore projects & source code | |
| jev-guardrailsgithub.com | 3 | Comparing LLM-as-judge vs TypeSafe Jev for agent guardrails: same rules, same agent, measured on cost, latency, calibration and coverage. | Benchmarks & evaluationMore projects & source code | |
| jev-guardrails — READMEgithub.com | 3 | A mock support agent benchmarks two guardrail backends—an LLM judge and typed Jev judgments—on the same rules; its reported measurements are project results. | Benchmarks & evaluationMore projects & source code | |
| jev-hankogithub.com | 0 | Measuring TypeSafe AI's Jev on 41-clause contract review (CUAD, 20,500 decisions) against fast, cheap LLMs — latency, cost and F1 | Benchmarks & evaluationMore projects & source code | |
| jev-hanko — READMEgithub.com | 0 | A benchmark compares Jev with other models on classifying contract pages against 41 clause types. | Benchmarks & evaluationMore projects & source code | |
| jev-headline-bench — READMEgithub.com | 0 | This benchmark asks whether Jev can choose the winner of a real headline A/B test; the reason here is the experiment design, not its reported score. | Benchmarks & evaluationMore projects & source code | |
| jev-hftgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-hft — READMEgithub.com | 0 | This research project examines Jev judgments over market data and news in a trading context; the README frames it as an evaluation, not a trading guarantee. | Benchmarks & evaluationMore projects & source code | |
| jev-integration-reportgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-integration-report — READMEgithub.com | 0 | A report documents an integration involving agy-opencode-jev; this repository is reporting material, not a reusable Jev client. | Benchmarks & evaluationMore projects & source code | |
| jev-jp-accounting-benchgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-jp-accounting-bench — READMEgithub.com | 0 | This repository adapts Japanese accounting tasks for Jev-compatible System One models; its README distinguishes the adaptation from official JMMLU/jfinqa evaluation and warns against direct score comparison. | Benchmarks & evaluationMore projects & source code | |
| jev-jp-addressgithub.com | 1 | CLI that normalizes messy Japanese addresses to postal codes against Japan Post's KEN_ALL master, using rules first and asking Jev a Choice over candidates only where matching is ambiguous. | Benchmarks & evaluationMore projects & source code | |
| jev-jp-address — READMEgithub.com | 1 | This Japanese-address CLI uses Jev Choice only when postal-master and rule matching cannot resolve a field, then lets code assemble the normalized address. | Benchmarks & evaluationMore projects & source code | |
| jev-judge-benchgithub.com | 0 | Binary LLM-judge bench — compare Jev (TypeSafe System One) against a frontier LLM judge on speed, cost, and agreement with human labels. | Benchmarks & evaluationMore projects & source code | |
| jev-korean-benchmark — READMEgithub.com | 6 | A frozen, reproducible 100-question-per-cell sample check compares Jev on Korean and English text; its authors caution that small differences are not findings. | Benchmarks & evaluationMore projects & source code | |
| jev-labgithub.com | 0 | Six small apps and a workbench that show what TypeSafe's Jev (System One) model can do. | Benchmarks & evaluationMore projects & source code | |
| jev-labgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-lab — READMEgithub.com | 0 | jev-lab is a TypeScript CLI for trying Jev on synthetic support text with urgency, department, and frustration questions; its README says tests use mocks. | Benchmarks & evaluationMore projects & source code | |
| jev-lab — READMEgithub.com | 1 | An independent lab explores fitting Jev beside Claude Code, including per-turn model routing and failure fallback; author-reported measurements are not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-lab — READMEgithub.com | 1 | An independent research lab curates Jev/System One resources, use cases, and reproducible benchmarks; it is an experiment and reference catalog. | Benchmarks & evaluationMore projects & source code | |
| jev-lab — READMEgithub.com | 0 | This browser-agent safety harness describes a generative model proposing actions and Jev selecting the next action with a separate typed safety judgment; it ships no benchmark numbers. | Benchmarks & evaluationMore projects & source code | |
| jev-lab — READMEgithub.com | 0 | A research repository records four independent experiments on TypeSafe Jev and links results to committed raw response logs; reported findings were not independently reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-labsgithub.com | 1 | TLA+-verified consensus kernel around Jev, generated into Rust and run through 1,680 simulated pharmacy decisions under seeded chaos against the live API, with zero wrong verdicts and more escalations as evidence degrades. | Benchmarks & evaluationMore projects & source code | |
| jev-lifegithub.com | 0 | The Chess of Life × Jev — an experimental game: write a new ruleset, then watch a decision model play it. | 生命棋 × Jev:实验性游戏设计——写一套新规则,然后看 Jev 怎么玩 | Benchmarks & evaluationMore projects & source code | |
| jev-life — READMEgithub.com | 0 | An experimental ruleset game uses Jev to explore and play each new board with typed legal-cell choices; the repo also describes a separate Jev-compatible LLM broker. | Benchmarks & evaluationMore projects & source code | |
| jev-little-airways — READMEgithub.com | 6 | A toy flight-simulation study assigns in-flight choices to live Jev calls, while the README attributes the visuals to Astra and simulation/ATC logic to GLM. | Benchmarks & evaluationMore projects & source code | |
| jev-llm-benchmarkgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-llm-benchmark — READMEgithub.com | 0 | A benchmark project compares TypeSafe Jev with a conventional LLM for latency and accuracy; this confirms the topic only, not reported results. | Benchmarks & evaluationMore projects & source code | |
| jev-llm-judgegithub.com | 0 | Can a decision model replace an LLM as an eval judge? Same rubrics, scored by Jev and by an LLM, side by side. | Benchmarks & evaluationMore projects & source code | |
| jev-llm-judge — READMEgithub.com | 0 | Evaluation research ports five agent-trace rubrics to Jev typed Score questions and compares them with a chat-model judge through a shared EvalResult contract, without generating Jev explanations. | Benchmarks & evaluationMore projects & source code | |
| Jev-LLM-Playground — READMEgithub.com | 0 | The community playground includes a Jev-based support-routing CLI, keyword baseline, and offline replay. | Benchmarks & evaluationMore projects & source code | |
| jev-lm — READMEgithub.com | 6 | A word-level drafting tool uses Jev to probe, verify and score candidate continuations; tokenization, sampling, repetition control and stopping remain ordinary code. | Benchmarks & evaluationMore projects & source code | |
| jev-mahjong-benchgithub.com | 1 | Reproducible riichi mahjong benchmark for Jev, GPT, Mortal, and hybrid agents using MJAI and RiichiEnv. | Benchmarks & evaluationMore projects & source code | |
| jev-mahjong-bench — READMEgithub.com | 1 | The README reports an author-run comparison of Jev and GPT for riichi-mahjong choices on matching game states; the results were not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-mario — READMEgithub.com | 1 | jev-mario is an emulator harness in which TypeSafe Jev selects Mario actions from RAM-derived game state; README distinguishes branching and direct-control modes. | Benchmarks & evaluationMore projects & source code | |
| jev-mcp — READMEgithub.com | 7 | An MCP server exposes TypeSafe Jev as a typed judgment tool and pairs it with an evaluation CLI that compares results against general-purpose LLMs. | Benchmarks & evaluationMore projects & source code | |
| jev-music-theory-1 — READMEgithub.com | 2 | Jev Chorale Lab describes a music-theory experiment in which Jev answers typed multiple-choice questions and a separate rule-based grader evaluates its work. | Benchmarks & evaluationMore projects & source code | |
| jev-nflgithub.com | 0 | Before every NFL snap, a decision-only AI model (TypeSafe Jev) calls run or pass and go/punt/kick on fourth down, graded live against the coach. | Benchmarks & evaluationMore projects & source code | |
| jev-nfl — READMEgithub.com | 0 | An NFL play-tracking project records Jev’s pre-snap run/pass and fourth-down choices, then compares them with the coach’s actions. | Benchmarks & evaluationMore projects & source code | |
| jev-no-enem — READMEgithub.com | 0 | Jev no ENEM is a preliminary evaluation of TypeSafe Jev on Brazilian standardized-exam questions; the README says more trials are needed for formal statistical significance. | Benchmarks & evaluationMore projects & source code | |
| jev-noul-vs-choicegithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-noul-vs-choice — READMEgithub.com | 0 | A research project compares Jev’s dice-related probability behavior across question formats, asking whether the effect is general or tied to how the question is asked. | Benchmarks & evaluationMore projects & source code | |
| jev-obniz-ledgithub.com | 0 | Jev × obniz LED: Physical AI Hello World — text → typed decisions (TypeSafe System One) → WS2812B LEDs | Benchmarks & evaluationMore projects & source code | |
| jev-obniz-led — READMEgithub.com | 0 | This Jev/obniz demo maps typed judgments from a user sentence to LED color, count, brightness, and blinking; code maps the values to the physical LEDs. | Benchmarks & evaluationMore projects & source code | |
| jev-on-a-laptopgithub.com | 24 | Unofficial study reproducing Jev-style parallel constrained decoding on stock Qwen 1.5B-8B models on an Apple Silicon laptop, where 7B reached 73.8% agreement against Jev's 86.6% at 0.4-2 s per decision. | Benchmarks & evaluationMore projects & source code | |
| jev-on-a-laptop — READMEgithub.com | 24 | This research repo reproduces the parallel-constrained-decoding idea behind Jev on a stock local model; it is a study, not a TypeSafe API compatibility claim. | Benchmarks & evaluationMore projects & source code | |
| jev-orderby-benchgithub.com | 0 | Measures whether ORDER BY over a Jev probability is defensible: pairwise inversion, Score ordinality against a human grade, calibration, and wording invariants under a pre-registered gate; passes on 20 Newsgroups topics, fails four of six conditions on Amazon ESCI product relevance, and shows that a 40-row batched state through a DuckDB extension fails the ranking gate one row per request passes. | Benchmarks & evaluationMore projects & source code | |
| jev-orderby-bench — READMEgithub.com | 0 | An independent README reports a preregistered evaluation of Jev probability ordering on 360 human-labeled rows; results are author-reported, not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-packsgithub.com | 0 | Registry of Jev question packs (questions, golden cases, pinned models, measured evidence) plus Jev Bench, which scores jev-1.13.0, claude-sonnet-5 and a local Qwen on the same labels with accuracy, ECE, cost and latency. | Benchmarks & evaluationMore projects & source code | |
| jev-packs — READMEgithub.com | 0 | This repository provides a registry and golden-set benchmark of Jev question packs; reported scores are the authors’ measurements, not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-phishing-benchgithub.com | 7 | Reproducible benchmark of Jev versus Claude Haiku 4.5 on whether an email agent should click the link in 2,000 emails, where Jev scored 62.6% accuracy to Haiku's 81.3%, with a calibration audit. | Benchmarks & evaluationMore projects & source code | |
| jev-phishing-bench — READMEgithub.com | 7 | This repository benchmarks whether an email agent should click an email link and identifies Jev as the TypeSafe System One model being examined. | Benchmarks & evaluationMore projects & source code | |
| jev-physical-ai — READMEgithub.com | 3 | This report applies TypeSafe Jev to robotics, fleet triage, and edge-device comparisons; its README says the incident data is simulated, not production failure accuracy. | Benchmarks & evaluationMore projects & source code | |
| jev-pick-and-place-study — READMEgithub.com | 1 | This MuJoCo pilot compares Jev pick-and-place choices with reactive rules and Claude; its README reports Jev selected the same action sequence as the rules in 20 episodes. | Benchmarks & evaluationMore projects & source code | |
| jev-playgroundgithub.com | 0 | Weekly builds on Jev (TypeSafe AI), benchmarked honestly enough to publish. Week 01: inbound lead triage, 90% routing at 366ms and $0.04 per thousand leads. | Benchmarks & evaluationMore projects & source code | |
| jev-playgroundgithub.com | 0 | Local JEV experiments: route decisions, Minesweeper solvers, and drone simulation | Benchmarks & evaluationMore projects & source code | |
| jev-playgroundgithub.com | 19 | Japanese MoonBit and TypeScript playground of Jev experiments: a MoonBit client, Jev-vs-Jev gomoku, headless 3v3 and 5v5 MOBAs, a shell-command risk gate, an ESLint plugin and a Jev-judgment language. | Benchmarks & evaluationMore projects & source code | |
| jev-playgroundgithub.com | 1 | Toy projects testing Jev (TypeSafe System One): snake autopilot, Todoist sorter, HN re-ranker, Dangerous Dave autopilot | Benchmarks & evaluationMore projects & source code | |
| jev-playground — READMEgithub.com | 1 | Jev Playground is a small Next.js app for experimenting with TypeSafe AI’s Jev model. | Benchmarks & evaluationMore projects & source code | |
| jev-playground — READMEgithub.com | 0 | Experiment collection explores Jev typed decisions in lead triage and deal-risk workflows, reporting project-authored measurements rather than independent benchmarks. | Benchmarks & evaluationMore projects & source code | |
| jev-playground — READMEgithub.com | 0 | The repository contains experiments; the route demo compares TypeSafe Jev and a conventional LLM on shared route context. | Benchmarks & evaluationMore projects & source code | |
| jev-playground — READMEgithub.com | 1 | The playground groups toy Jev experiments, including Snake direction choices and sorting a Todoist inbox. | Benchmarks & evaluationMore projects & source code | |
| jev-playground — READMEgithub.com | 2 | A music-composition playground uses Jev, or an offline stub, to choose bounded labels; application code renders notes and exports MIDI. | Benchmarks & evaluationMore projects & source code | |
| jev-playsgithub.com | 3 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-plays — READMEgithub.com | 3 | A Craftax survival-game experiment pairs Jev's move selection with GPT-5.6-terra goal planning; it describes a game simulation, not physical autonomy. | Benchmarks & evaluationMore projects & source code | |
| jev-pocgithub.com | 0 | TypeSafe AI の判定モデル jev に 2048 を遊ばせる PoC(Go CLI + Cloudflare Workers の Web デモ) | Benchmarks & evaluationMore projects & source code | |
| jev-poc — READMEgithub.com | 0 | A Jev 2048 proof of concept records failed API games instead of filling missing moves with random choices; reported game scores were not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-practice-speed — READMEgithub.com | 1 | A WebGL Speed card-game demo pits the player against Jev and displays decision-speed and accuracy indicators; a mechanical validator checks whether moves are legal. | Benchmarks & evaluationMore projects & source code | |
| jev-prior-auth-triagegithub.com | 0 | Prior-authorization triage using TypeSafe AI's Jev (System-1 model) — payer-side utilization management, synthetic PHI-free data, audit-logged decisions. | Benchmarks & evaluationMore projects & source code | |
| jev-prior-auth-triage — READMEgithub.com | 0 | The README presents Jev prior-authorization triage routes to auto-approval, human review, or peer escalation rather than denial. | Benchmarks & evaluationMore projects & source code | |
| jev-probe — READMEgithub.com | 0 | Exploration harness compares Jev and baseline LLM decision consistency under meaning-preserving input perturbations and records confidence, cost and latency metrics. | Benchmarks & evaluationMore projects & source code | |
| jev-project-fit-reviewgithub.com | 1 | Review independente sobre a adequação do Jev a produtos e desenvolvimento multiagente | Benchmarks & evaluationMore projects & source code | |
| jev-project-fit-review — READMEgithub.com | 1 | A Portuguese independent design review evaluates where Jev can assist product and multi-agent decisions while deterministic rules retain authority; it is not a deployed integration. | Benchmarks & evaluationMore projects & source code | |
| jev-projectsgithub.com | 0 | Small demos of Jev (TypeSafe) through the Vercel AI Gateway: wiki race, town of agents, bullet chess, and more | Benchmarks & evaluationMore projects & source code | |
| jev-projects — READMEgithub.com | 0 | The repository collects Jev project demos, including a ViZDoom run, with shared Jev calls in lib.mjs. | Benchmarks & evaluationMore projects & source code | |
| jev-prompt-optimizationgithub.com | 1 | automatically optimizing the instructions and decision criteria of TypeSafe Jev Choice from labeled data | Benchmarks & evaluationMore projects & source code | |
| jev-prompt-optimization — READMEgithub.com | 1 | A labeled-data optimizer searches TypeSafe Jev Choice instructions and criteria using an evolutionary process with GPT as a mutation operator. | Benchmarks & evaluationMore projects & source code | |
| jev-quiz-pilot — READMEgithub.com | 0 | Browser quiz pilot reads questions and options, asks Jev to choose answers and clicks them, providing a harness for observing quiz performance. | Benchmarks & evaluationMore projects & source code | |
| jev-rag-benchmarkgithub.com | 14 | Reproducible benchmark of Jev as the reranker in a small RAG system on 1,044 Turkish XQuAD questions, measuring quality, latency and cost with the same 20 candidates given to every reranker. | Benchmarks & evaluationMore projects & source code | |
| jev-rag-benchmark — READMEgithub.com | 1 | README presents a RAG benchmark comparing Jev 1.13 with an NVIDIA cross-encoder and a no-reranking baseline; results are not independently reproduced. | Benchmarks & evaluationMore projects & source code | |
| jev-ragcheckgithub.com | 0 | RAG evaluation with typed decisions: sentence-level hallucination verdicts with offsets and citations, passage and answer relevance, in one Jev call per answer. Benchmarked against Ragas, DeepEval, an LLM judge and HHEM. | Benchmarks & evaluationMore projects & source code | |
| jev-ragcheck — READMEgithub.com | 0 | RAG evaluator splits answers in code and uses Jev fixed-option judgments to associate sentence-level verdicts with retrieved evidence. | Benchmarks & evaluationMore projects & source code | |
| jev-reflex-autonomy-lab — READMEgithub.com | 19 | An interactive multi-drone autonomy simulation explores Jev’s fast System 1 judgments with optional slower System 2 guidance. | Benchmarks & evaluationMore projects & source code | |
| jev-replay-lab — READMEgithub.com | 0 | Jev Replay Lab is a local workbench for labeled datasets, threshold simulation, error inspection, and run comparison; its synthetic demo is not evidence of Jev accuracy. | Benchmarks & evaluationMore projects & source code | |
| jev-reportgithub.com | 1 | Independent Chinese research report on Jev: a 52-page PDF that traces vendor numbers to their sources, plus 50 Chinese test cases where 86 of 90 judgments were correct (95.6%) with ECE 0.070. | Benchmarks & evaluationMore projects & source code | |
| jev-report — READMEgithub.com | 1 | This is an independent Chinese-language report about TypeSafe Jev, with a reproducible 50-sample Chinese evaluation described in the README. | Benchmarks & evaluationMore projects & source code | |
| jev-rerank-bench — READMEgithub.com | 9 | A saved reranking benchmark compares Jev with Cohere and ZeroEntropy; its README reports near-tied averages without establishing a winner, not independently reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-researchgithub.com | 0 | Jev-driven research engine for coding agents: query understanding, RRF reranking, shadow next-best-source, benchmarks | Benchmarks & evaluationMore projects & source code | |
| jev-research — READMEgithub.com | 0 | jev-research uses Jev’s typed judgments for retrieval decisions such as candidate ranking and whether to continue searching. | Benchmarks & evaluationMore projects & source code | |
| jev-research-eval — READMEgithub.com | 2 | A reproducible research harness evaluates a pinned Jev Ultrafast research-browser session; it drives upstream rather than forking Jev. | Benchmarks & evaluationMore projects & source code | |
| jev-research-pipeline — READMEgithub.com | 3 | A daily research monitor where code owns collection loops, Jev judges relevance and an independently chosen LLM writes. Included as a pilot workflow, not proven research accuracy. | Benchmarks & evaluationMore projects & source code | |
| jev-rl — READMEgithub.com | 3 | JevRL is a reward-model lab using Jev scores in four classic game environments with DQN agents and saved replays; outcome claims remain the repository’s experiments. | Benchmarks & evaluationMore projects & source code | |
| jev-roastgithub.com | 0 | Score declared writing dimensions and cite only exact source spans for weak results. | Benchmarks & evaluationMore projects & source code | |
| jev-roast — READMEgithub.com | 0 | jev-roast uses Jev for rubric scores and source-span selection, while the README says it does not generate replacement copy. | Benchmarks & evaluationMore projects & source code | |
| jev-routing-experiment — READMEgithub.com | 2 | This RouterArena study tests Jev as an LLM router; its README says the no-Jev ablation matched the reported result. | Benchmarks & evaluationMore projects & source code | |
| jev-sandbox — READMEgithub.com | 1 | This research sandbox tests Jev on email triage and ingredient-list gluten detection; the README frames these as experiments, not independently reproduced results. | Benchmarks & evaluationMore projects & source code | |
| jev-screengithub.com | 0 | Title and abstract screening for systematic reviews with Jev: explicit criteria, include/exclude/maybe with reasons, PRISMA counts, RIS export, recall against human screeners. | Benchmarks & evaluationMore projects & source code | |
| jev-screen — READMEgithub.com | 0 | jev-screen supports title/abstract screening against explicit review criteria, producing include/exclude/maybe buckets and human-review paths. Its README calls it a second screener or prioritizer, never a sole decider. | Benchmarks & evaluationMore projects & source code | |
| jev-sec-bench — READMEgithub.com | 3 | A repository of blind security benchmarks for Jev's typed judgments; a security evaluation suite is not itself evidence of improved security. | Benchmarks & evaluationMore projects & source code | |
| jev-secret-detectiongithub.com | 2 | Benchmark of how well Jev spots real, usable secret credentials in 100 file snippets plus edge and config sets, reporting accuracy, AUC, and recall with no regex or provider verification. | Benchmarks & evaluationMore projects & source code | |
| jev-secret-detection — READMEgithub.com | 2 | The README describes an evaluation that asks Jev whether file snippets contain usable secret credentials; it is a benchmark, not secret-scanning proof. | Benchmarks & evaluationMore projects & source code | |
| jev-serp-opportunity-labgithub.com | 0 | Terminal-based Jev intent triage with typed decisions, explicit simulation mode, review policy and JSON exports. | Benchmarks & evaluationMore projects & source code | |
| jev-serp-opportunity-lab — READMEgithub.com | 0 | Search-intent triage separates Jev judgments from code routing and includes an API connector plus offline policy simulator; bundled SERP examples are manufactured fixtures. | Benchmarks & evaluationMore projects & source code | |
| jev-shadcn-lint-evalgithub.com | 0 | Second eval for shadcn-ui/lint that asks Jev whether each of 131 rule test cases is right and whether its message says what to change; Jev caught 91% of real violations but was weaker on clean code. | Benchmarks & evaluationMore projects & source code | |
| jev-shadcn-lint-eval — READMEgithub.com | 0 | This evaluation repository asks Jev whether shadcn/ui lint findings are correct and whether their messages explain what to change; it is an evaluation, not a linter replacement. | Benchmarks & evaluationMore projects & source code | |
| jev-shadowgithub.com | 0 | Test whether TypeSafe Jev answers your questions correctly before you let it decide anything. | Benchmarks & evaluationMore projects & source code | |
| jev-shadow — READMEgithub.com | 0 | jev-shadow compares Jev’s judgments on previously decided support tickets with the user’s labels; its demo uses fake answers, not Jev output. | Benchmarks & evaluationMore projects & source code | |
| jev-shortlistgithub.com | 0 | A cached Jev prior for active learning: reusable rank fusion, matched ASReview controls, and a no-key evidence replay. | Benchmarks & evaluationMore projects & source code | |
| jev-shortlist — READMEgithub.com | 0 | A research tool caches Jev judgments and combines rankings for active learning; its README includes saved benchmark traces, not rerun in this review. | Benchmarks & evaluationMore projects & source code | |
| jev-showdowngithub.com | 0 | Jev plays Pokemon Showdown: 40 battles, 940 decisions, reproducible results, and an annotated replay. | Benchmarks & evaluationMore projects & source code | |
| jev-showdown — READMEgithub.com | 0 | This local Pokémon Showdown experiment gives Jev every legal move and switch, then saves the game state it saw and the choice it made. | Benchmarks & evaluationMore projects & source code | |
| jev-simgithub.com | 1 | Jev-compatible /v1/systemone server reading typed decisions from LLM logits, benchmarked against TypeSafe's Jev on the same items via JevBench | Benchmarks & evaluationMore projects & source code | |
| jev-sim — READMEgithub.com | 1 | The README describes Jev-sim as a wire-compatible reimplementation that reads typed decisions from LLM logits and includes a comparison with Jev; compatibility was not tested here. | Benchmarks & evaluationMore projects & source code | |
| jev-snake — READMEgithub.com | 0 | This browser-based Snake benchmark races TypeSafe Jev against a frontier LLM under the same board seed and time budget. | Benchmarks & evaluationMore projects & source code | |
| jev-snap-labgithub.com | 0 | Tiny inputs. Instant decisions. A small experimental playground for exploring fast, probabilistic decisions with Jev. | Benchmarks & evaluationMore projects & source code | |
| jev-snap-lab — READMEgithub.com | 0 | Jev Snap Lab is a short-text web app that visualizes multiple small Jev judgments; its CITY mode is described as a technical demo, not official administrative guidance. | Benchmarks & evaluationMore projects & source code | |
| jev-spam-eval — READMEgithub.com | 2 | This repository applies Jev to email-category decisions: code supplies an email and category definitions, and Jev returns category probabilities. | Benchmarks & evaluationMore projects & source code | |
| jev-storyboard-labgithub.com | 3 | Google ADK vs Microsoft Agent Framework for structured-output agents, with TypeSafe Jev as a vendor-neutral QC gate | Benchmarks & evaluationMore projects & source code | |
| jev-studygithub.com | 0 | Hands-on measurements of TypeSafe's Jev via OpenRouter: integration, Chinese-language behaviour, and game loops. Every figure in the reports is reproducible by the scripts here. Reports in Chinese. ~1,430 API calls, ~$0.15. | Benchmarks & evaluationMore projects & source code | |
| jev-study — READMEgithub.com | 0 | OpenRouter-based Jev study publishes scripts and Chinese reports on API integration, language behavior and game loops, including chess/xiangqi; its tests cover a pinned model snapshot rather than every provider. | Benchmarks & evaluationMore projects & source code | |
| jev-support-pulse — repogithub.com | 1 | Chinese experiment labeling 170,400 2017 tweets to seven brands' support accounts with Jev for $1.84: at equal false alarms it caught 17 outages about 4.1 hours before the brand admitted them, versus 10 for tweet volume. | Benchmarks & evaluationMore projects & source code | |
| jev-synergy-screening — READMEgithub.com | 2 | The ADHD abstract-triage demo uses TypeSafe Jev for MEDLINE title-and-abstract include/exclude judgments with Choice and Noul questions. | Benchmarks & evaluationMore projects & source code | |
| jev-synthetic-survey — READMEgithub.com | 2 | This independent survey study compares TypeSafe Jev’s native probability outputs with verbalized estimates on synthetic respondents, without treating the README results as a general benchmark. | Benchmarks & evaluationMore projects & source code | |
| jev-system-onegithub.com | 0 | A polished OpenAI + TypeSafe Jev terminal interface for answers with transparent decision reports. | Benchmarks & evaluationMore projects & source code | |
| jev-testgithub.com | 0 | Playground for TypeSafe's Jev System One model (Next.js) | Benchmarks & evaluationMore projects & source code | |
| jev-testgithub.com | 0 | Reproducible zero-shot Jev benchmark on all seven LexGLUE tasks | Benchmarks & evaluationMore projects & source code | |
| jev-testgithub.com | 1 | Pre-registered benchmark: can a 2B local model (Gemma 4 E2B) answer web questions without making things up when a decision model (TypeSafe Jev) makes every call? SearXNG for search, MemPalace for verbatim memory, seven arms including open local judges. Spec and thresholds fixed before any run. | Benchmarks & evaluationMore projects & source code | |
| jev-testgithub.com | 0 | Pruebita usando Jev: AI slop detector en Twitter | Benchmarks & evaluationMore projects & source code | |
| jev-testgithub.com | 0 | Empirical benchmarks, task telemetry, and engineering thesis for Jev (TypeSafe AI) System One in autonomous agentic organizations | Benchmarks & evaluationMore projects & source code | |
| jev-test — READMEgithub.com | 0 | This small Next.js playground visualizes Jev’s typed answers and raw request/response; the README describes a demo rather than a verified compatibility test. | Benchmarks & evaluationMore projects & source code | |
| jev-test — READMEgithub.com | 0 | A reproducible zero-shot Jev/LexGLUE benchmark sends typed classification questions through OpenRouter; reported scores are documentation claims, not independently reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-test — READMEgithub.com | 1 | jev-test is a research harness using TypeSafe Jev as a separate judge for search, evidence selection, and sentence support; its README also reports the judged pipeline missed its baseline. | Benchmarks & evaluationMore projects & source code | |
| jev-test — READMEgithub.com | 0 | A research prototype includes a TypeSafe SDK smoke test and a Chrome extension that flags X/Twitter posts Jev judges to be low-effort AI-generated content; no accuracy claim is verified here. | Benchmarks & evaluationMore projects & source code | |
| jev-test — READMEgithub.com | 0 | This engineering-research repository describes a Jev/System One architecture thesis and reports several benchmarks; results are the authors’ claims and were not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-testbedgithub.com | 0 | Jev (TypeSafe System One) 테스트베드 — 클라우드 API와 로컬 셀프호스팅(jeff/GLiFormer) 양쪽 실행 예제 및 실측 결과 | Benchmarks & evaluationMore projects & source code | |
| jev-testbed — READMEgithub.com | 0 | A Korean-language testbed explains Jev’s typed decision outputs and states that Jev itself is hosted-only. | Benchmarks & evaluationMore projects & source code | |
| jev-tetris-benchmark — READMEgithub.com | 0 | The README describes a Tetris comparison that gives Jev and Claude Haiku the same legal placement candidates; benchmark results were not independently reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-tick-labgithub.com | 0 | A forward-only experiment: Jev (TypeSafe System One) making one-second trading judgments on bitbank, logged for calibration analysis. | Benchmarks & evaluationMore projects & source code | |
| jev-tool-search — READMEgithub.com | 1 | A tool-selection benchmark compares Jev with lexical, embedding and reranking approaches, alongside an experimental Jev search engine. | Benchmarks & evaluationMore projects & source code | |
| jev-trace-classifier — READMEgithub.com | 0 | The benchmark applies TypeSafe Jev to the public collusion.wiki corpus to classify whether pages were written by an autonomous agent or a human. | Benchmarks & evaluationMore projects & source code | |
| jev-trade — READMEgithub.com | 1 | jev-trade is a simulated crypto-trading loop that uses Jev for typed buy/sell decisions; it describes simulated execution, not live trading. | Benchmarks & evaluationMore projects & source code | |
| jev-traffic-racegithub.com | 0 | Live demo showing why loop-speed classification matters: Jev vs LLMs on the same events, same clock, honest scorecard. | Benchmarks & evaluationMore projects & source code | |
| jev-traffic-race — READMEgithub.com | 0 | Jev Reflex is a two-lane demo comparing Jev and a frontier LLM on the same Bombay road-hazard classification task. This review does not treat its visualization as a formal benchmark or infer general superiority. | Benchmarks & evaluationMore projects & source code | |
| jev-translation-checkergithub.com | 0 | Experimental PoC for Jev translation checking: bilingual benchmarks, prompt comparisons, re-verification, and MAGI voting. | Benchmarks & evaluationMore projects & source code | |
| jev-translation-checker — READMEgithub.com | 0 | An experimental proof of concept uses Jev to check translation error patterns and guide a bounded repair loop for a separate translation LLM; the README makes no production-quality guarantee. | Benchmarks & evaluationMore projects & source code | |
| jev-typesafe-real-financial-use-cases — READMEgithub.com | 0 | Jev Lab is a financial-use-case demo and backtest site built on TypeSafe’s JavaScript SDK; the README says live mode is local and opt-in. | Benchmarks & evaluationMore projects & source code | |
| jev-typesafe-spikegithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jev-typesafe-spike — READMEgithub.com | 0 | Experimental data-extraction workflow uses DeepSeek to extract information from text and Jev for subsequent auditing; the earlier structure-detection plan is suspended. | Benchmarks & evaluationMore projects & source code | |
| jev-typesafeai-testgithub.com | 0 | Testing TypesafeAI's "Jev" model using System One API | Benchmarks & evaluationMore projects & source code | |
| jev-typesafeai-test — READMEgithub.com | 0 | This small test of TypeSafe System One sends typed questions about a support ticket to Jev. | Benchmarks & evaluationMore projects & source code | |
| Jev-Visiongithub.com | 4 | Open-weight step verifier for computer-use agents: calibrated ground/skip/effect/done judgments from screenshots in ~160 ms, plus a benchmark with environment-derived labels | Benchmarks & evaluationMore projects & source code | |
| Jev-Vision — READMEgithub.com | 4 | Jev-Vision describes a separate 8B vision model using TypeSafe Jev’s typed request/response shape, not the hosted Jev model. | Benchmarks & evaluationMore projects & source code | |
| jev-vs-llm-stock-policy — READMEgithub.com | 0 | This policy demo describes Jev’s typed Choice/Score/Noul outputs; its no-key SAMPLE values are author-filled rather than Jev measurements. | Benchmarks & evaluationMore projects & source code | |
| jev-vs-llm-ticket-routergithub.com | 0 | Benchmark: TypeSafe Jev vs traditional LLM on support-ticket routing accuracy, latency, and cost | Benchmarks & evaluationMore projects & source code | |
| jev-vs-llm-ticket-router — READMEgithub.com | 0 | Python support-ticket benchmark compares Jev typed classification against an OpenRouter LLM using the same department labels, tracking latency and estimated cost. | Benchmarks & evaluationMore projects & source code | |
| jev-vs-luna — READMEgithub.com | 0 | This benchmark compares Jev through OpenRouter’s Decisions API with Luna through chat completions on the same review text; it is a workload-specific comparison. | Benchmarks & evaluationMore projects & source code | |
| jev-vs-tfidf-benchmark — READMEgithub.com | 0 | This reproducible benchmark compares Jev with TF-IDF/Jaccard matching for disguised duplicate thesis titles; the repository’s reported measurements are not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| jev-wavegithub.com | 0 | TypeSafe Jev research + 5 product specs. | Benchmarks & evaluationMore projects & source code | |
| jev-writergithub.com | 2 | Find out which qualities of your writing actually predict engagement. Rates every post you have published against a pre-registered rubric using Jev's calibrated judgments, then tests those ratings against your real engagement numbers. Refuses to report findings your sample cannot support. | Benchmarks & evaluationMore projects & source code | |
| jev-writer — READMEgithub.com | 2 | jev-writer uses TypeSafe Jev to rate published posts against registered writing rubrics and compare those ratings with engagement outcomes. | Benchmarks & evaluationMore projects & source code | |
| jev-x-postsgithub.com | 0 | Sortable 48-hour X post report for Jev, TypeSafe AI and Diogo Almeida, with Jev sentiment labels. | Benchmarks & evaluationMore projects & source code | |
| jev_experimentsgithub.com | 0 | Experiments with TypeSafe's Jev System One model | Benchmarks & evaluationMore projects & source code | |
| jev_experiments — READMEgithub.com | 0 | This case-based repository documents experiments with TypeSafe Jev, including a Chat Noir demo where Jev chooses a candidate wall; the README does not establish general performance. | Benchmarks & evaluationMore projects & source code | |
| jev_playground — READMEgithub.com | 1 | An evidence-focused Jev evaluation playground separates recorded-answer demos from optional --live API calls. | Benchmarks & evaluationMore projects & source code | |
| jev_practicegithub.com | 0 | Jev (TypeSafe System One) と LLM に同じゲームを打たせて、レイテンシ・コスト・判断の質を比べる練習台 | Benchmarks & evaluationMore projects & source code | |
| jev_practice — READMEgithub.com | 0 | A Tetris practice project runs Jev alongside Claude Haiku 4.5; displayed scores and timing are author-reported, not independently tested. | Benchmarks & evaluationMore projects & source code | |
| jev_stock — READMEgithub.com | 14 | A Hong Kong stock-direction experiment uses TypeSafe Jev for up/flat/down forecasts and says it is not a proven trading model. | Benchmarks & evaluationMore projects & source code | |
| jev_typesafeai_testgithub.com | 1 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jeval — presets.pygithub.com | 21 | jeval documents a jev-native preset for recorded decision-API responses such as jev-1.13.0; it analyzes stored data rather than calling a model. | Benchmarks & evaluationMore projects & source code | |
| jeval — READMEgithub.com | 21 | jeval documents a jev-native preset for recorded decision-API responses such as jev-1.13.0; it analyzes stored data rather than calling a model. | Benchmarks & evaluationMore projects & source code | |
| jeval — READMEgithub.com | 0 | jeval is an independent evaluation SDK/CLI for AI outputs and agents, with TypeSafe Jev as its first judge provider. The README marks its cloud offering as planned, not built. | Benchmarks & evaluationMore projects & source code | |
| Jevalsgithub.com | 1 | Independent benchmark data comparing Jev with six LLMs on PubMedQA, Banking77 and HelpSteer2 against human labels, covering accuracy, calibration, cost and latency. | Benchmarks & evaluationMore projects & source code | |
| Jevals — READMEgithub.com | 0 | Jevals is an independent benchmark site that compares Jev with LLMs on shared typed questions graded against human labels; this review did not reproduce its reported findings. | Benchmarks & evaluationMore projects & source code | |
| jevals — READMEgithub.com | 99 | Jevals evaluates agent traces with Jev through documented TypeSafe/Vercel routes and also lists local Kev/Laya or chat-model backends. | Benchmarks & evaluationMore projects & source code | |
| jevals-datagithub.com | 1 | Independent benchmark data for TypeSafe's Jev (System One model) vs LLMs: accuracy, calibration, cost. Boards + per-decision logs, CC-BY-4.0 | Benchmarks & evaluationMore projects & source code | |
| jevals-data — READMEgithub.com | 1 | This repository publishes independent benchmark data for Jev and Jev-type models across Noul, Choice, and Score tasks. | Benchmarks & evaluationMore projects & source code | |
| JEValuategithub.com | 0 | Auto-marking maths scripts with Jev (TypeSafe System One): 2,054 scripts, 96.6% agreement with human markers | Benchmarks & evaluationMore projects & source code | |
| JEValuate — READMEgithub.com | 0 | JEValuate uses Jev with a teacher-style rubric to mark mathematics scripts and compares the results with human annotators. | Benchmarks & evaluationMore projects & source code | |
| jevalyzergithub.com | 2 | Grade the agent sessions already on your disk. Claude Code, Codex, opencode, Gemini CLI and Antigravity, scored with Jev for cents. | Benchmarks & evaluationMore projects & source code | |
| jevalyzer — READMEgithub.com | 2 | Jevalyzer reads local coding-agent chat logs and scores exchanges with TypeSafe Jev through the Vercel AI Gateway. | Benchmarks & evaluationMore projects & source code | |
| JevArenagithub.com | 4 | Hosted arena that puts Jev against an opponent judge you connect and takes your vote before revealing which was which, with latency, cost provenance, and self-reported confidence; a vote records preference, not verified correctness. | Benchmarks & evaluationMore projects & source code | |
| JevBenchgithub.com | 192 | Third-party benchmark that runs 534 frozen decision questions and folds intelligence, calibration, speed, and cost into one score across Jev, open reimplementations, classifiers, and LLM baselines. | Benchmarks & evaluationMore projects & source code | |
| JevBench (Benchmark Heaven)github.com | — | Benchmark for Jev and Jev-like decision models inside Benchmark Heaven, an LLM price and benchmark comparison site, with public and held-out task splits reported separately. | Benchmarks & evaluationMore projects & source code | |
| JevBench (Benchmark Heaven) — repogithub.com | — | Benchmark for Jev and Jev-like decision models inside Benchmark Heaven, an LLM price and benchmark comparison site, with public and held-out task splits reported separately. | Benchmarks & evaluationMore projects & source code | |
| jevbench — READMEgithub.com | 1 | An independent preregistered ChaosNLI study of whether Jev confidence falls when human annotators disagree. Included as empirical research, not proof of a calibration result. | Benchmarks & evaluationMore projects & source code | |
| jevbench — READMEgithub.com | 192 | JevBench is a benchmark for Jev-class typed-decision systems and compares Jev with other entries; its reported scores are repository claims, not independently rerun here. | Benchmarks & evaluationMore projects & source code | |
| JEVBenchmark-Contradiction-Detectiongithub.com | 0 | Checking JEV's Contradiction detection (Model by TypeSafe.AI). | Benchmarks & evaluationMore projects & source code | |
| jevbettergithub.com | 15 | One-pass option scorer built from scratch with a hashed n-gram encoder, rival-aware attention, a gated head and temperature scaling, trained on the jevlike data format and benchmarked head-to-head against it. | Benchmarks & evaluationMore projects & source code | |
| jevbetter — READMEgithub.com | 15 | An independent Jev-like one-pass option scorer reuses the jevlike dataset format; README explicitly says it is not affiliated with TypeSafe or Jev. | Benchmarks & evaluationMore projects & source code | |
| jevcalgithub.com | 10 | Toolkit that measures a typed decision model like Jev on your own labeled data against an LLM teacher, picks the confidence threshold for a target accuracy, reports how much traffic still needs an LLM, and fails CI on drift. | Benchmarks & evaluationMore projects & source code | |
| jevcheck — READMEgithub.com | 2 | A fixture-based contract-testing tool compares a candidate Jev model against pinned expected behavior; unit tests mock network calls. | Benchmarks & evaluationMore projects & source code | |
| JEVfiregithub.com | 71 | Jev-inspired parallel decisions for CUDA LLMs on vLLM that score single-token labels with the model's own head and assemble JSON in code, with game demos including in-browser Mario at 71 ms per action. | Benchmarks & evaluationMore projects & source code | |
| jevfire — READMEgithub.com | 71 | A local vLLM-based typed-choice scorer is inspired by JEV/RLCD and explicitly named in tribute; it is not a TypeSafe Jev client. | Benchmarks & evaluationMore projects & source code | |
| jevincigithub.com | — | Creative experiment: paints images by having Jev predict every pixel's colour in parallel, with predicted confidence deciding how wide each stroke is drawn. | Benchmarks & evaluationMore projects & source code | |
| jevlike — READMEgithub.com | 1,338 | Jevlike is described as an independent starter model with Jev’s input/output shape, making it a relevant research alternative rather than a TypeSafe model or tested drop-in. | Benchmarks & evaluationMore projects & source code | |
| jevllmgithub.com | 1 | Jev as an LLM (cz why not) | Benchmarks & evaluationMore projects & source code | |
| jevllm — READMEgithub.com | 1 | An experimental autoregressive text decoder repeatedly asks Jev to choose the next word; the repository describes it as a measurement/demo, not Jev’s native text generation. | Benchmarks & evaluationMore projects & source code | |
| jevmazegithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| jevmaze — READMEgithub.com | 0 | Jev Maze is a simulation that asks Jev to choose an up/down/left/right move one step at a time. | Benchmarks & evaluationMore projects & source code | |
| JevNextgithub.com | 5 | More than Choice — A simple algorithm that equips any Jev-like model with numerical control. | Benchmarks & evaluationMore projects & source code | |
| JevNext — READMEgithub.com | 5 | An experimental algorithm using successive Jev Choice branches to decode finite-precision numerical values. It does not add a native regression head or retrain Jev; calibration remains unverified. | Benchmarks & evaluationMore projects & source code | |
| jevoraclegithub.com | 0 | Typed questions over your own state, not a chat. A front end for TypeSafe's System One API. | Benchmarks & evaluationMore projects & source code | |
| jevoracle — READMEgithub.com | 0 | Prediction-market research dashboard records bot judgments before real market outcomes using Jev and optional local engines, but all bankrolls and bets are simulated rather than real-money transactions. | Benchmarks & evaluationMore projects & source code | |
| JevRankergithub.com | 5 | Jev-style decision models as BlitzRank's compare oracle: k passages scored in one parallel forward pass, zero decoded tokens — 22x faster per match than a generative listwise LLM, at 2.9x its nDCG@10. | Benchmarks & evaluationMore projects & source code | |
| JevRanker — READMEgithub.com | 5 | JevBlitzRank describes using Jev Choice in BlitzRank’s tournament in place of generative LLM calls. | Benchmarks & evaluationMore projects & source code | |
| Jevs-Garagegithub.com | 1 | A garage full of tiny experiments for building critical systems with System One & Jev 🔧🧠⚡ | Benchmarks & evaluationMore projects & source code | |
| Jevs-Garage — READMEgithub.com | 1 | The experiments give Jev typed questions over realistic state, while ordinary Python policy decides what happens next. | Benchmarks & evaluationMore projects & source code | |
| jevscape — READMEgithub.com | 12 | A RuneBench-derived RuneScape harness offers bounded game actions for Jev to choose; local code executes each action, so this is a specific game integration rather than all of RuneBench. | Benchmarks & evaluationMore projects & source code | |
| JevScopegithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| JevScope — READMEgithub.com | 0 | JevScope evaluates agent traces for semantic observability and robustness; its README describes it as an observer, not an execution gate. | Benchmarks & evaluationMore projects & source code | |
| jevtrafficsim — READMEgithub.com | 1 | This traffic simulation compares citywide policies using Jev; the README says Jev does not set individual signal states. | Benchmarks & evaluationMore projects & source code | |
| jevxgithub.com | 0 | Jev (TypeSafe System One) research: API notes, benchmarks, community experiments, agent-loop patterns | Benchmarks & evaluationMore projects & source code | |
| jevx — READMEgithub.com | 0 | A research collection and Python SDK use Jev (System 1) to drive Codex (System 2), according to the README. | Benchmarks & evaluationMore projects & source code | |
| jevygithub.com | 0 | Jev-style typed-decision model distilled from official Jev. 118M, EN+CN, trains on a 4GB GPU in 10 minutes. | Benchmarks & evaluationMore projects & source code | |
| jevy — READMEgithub.com | 0 | jevy is presented as a separately fine-tuned Jev-style typed-decision model distilled from the official Jev API, not the hosted TypeSafe model. | Benchmarks & evaluationMore projects & source code | |
| jitllm-jev-demogithub.com | 0 | System One + System Two on one GPU in pure Java: Jev-style decisions with jitLLM + TornadoVM | Benchmarks & evaluationMore projects & source code | |
| jitllm-jev-demo — READMEgithub.com | 0 | Java help-desk demo points a Jev API-contract client at local jitLLM for typed triage, then uses local generation for routed replies; it does not run hosted Jev. | Benchmarks & evaluationMore projects & source code | |
| judge-jevgithub.com | 0 | Reproducible experiments evaluating Jev as an automated judge across benchmarks and tasks. | Benchmarks & evaluationMore projects & source code | |
| judge-jev — READMEgithub.com | 0 | Research evaluates Jev Choice/Score/Noul as passage-relevance judges against human labels and published LLMJudge submissions; the public-label comparison is retrospective, not a blind challenge. | Benchmarks & evaluationMore projects & source code | |
| Judgment arena — repogithub.com | — | Video that explains System 1 vs System 2 and races Jev against Claude Opus, Haiku 4.5 and GPT-5.4 Mini on 15 human-labelled questions; Opus got one more right but took 10x longer and cost 146x more. | Benchmarks & evaluationMore projects & source code | |
| krino — README_engithub.com | 0 | Decision-model research combines black-box Jev API probes, open model replication, benchmark/data pipelines and architecture analysis; reported replication scores are research results, not a compatibility certification. | Benchmarks & evaluationMore projects & source code | |
| KsanaDock/verdict-labgithub.com | 1 | An experiment comparing the capabilities and costs of the JEV model and LLM models in the field of content moderation | Benchmarks & evaluationMore projects & source code | |
| kyotsu-ai-bench — en.htmlgithub.com | 1 | A Japanese Common Test comparison page reports Jev alongside four other model configurations across 23 subjects and 836 questions; these are the project’s published results. | Benchmarks & evaluationMore projects & source code | |
| LegalForecastBench Jev mode — repogithub.com | 5 | Official Jev condition in a benchmark that forecasts federal motion-to-dismiss rulings: Jev gets one case record, full text or LLM summaries, and is scored with claim-defendant micro-Brier metrics. | Benchmarks & evaluationMore projects & source code | |
| let-jev-speak — READMEgithub.com | 2 | This experiment generates prose by decoding one word at a time through successive Jev choice questions, rather than a native text-generation call. | Benchmarks & evaluationMore projects & source code | |
| LightJevgithub.com | 1 | Train lightweight language backbones for typed decisions and candidate probabilities. CE/Brier training, evaluation, and an offline end-to-end demo. | Benchmarks & evaluationMore projects & source code | |
| Little Airwaysgithub.com | 6 | Toy archipelago flight simulator where each aircraft sees only itself and nearby traffic and Jev decides live whether to divert, declare an emergency, give way, hold or land first, in about 150 ms. | Benchmarks & evaluationMore projects & source code | |
| LLM Chess: Jev resultsgithub.com | 131 | Long-running chess benchmark for LLMs that added Jev: across 80 games against a random player and Komodo Dragon it made zero illegal moves, winning 8 and drawing 22. | Benchmarks & evaluationMore projects & source code | |
| mallahyari/system-one-benchmarkgithub.com | 1 | System One & Parallel Constrained Decoding Benchmark | Benchmarks & evaluationMore projects & source code | |
| Minutes Jev voice evalsgithub.com | — | Synthetic qualification script in the Minutes meeting-memory app that tests Jev on seven Choice decisions its voice path needs, such as attendee constraints, semantic recall, verified pastes, stale targets and prompt injection. | Benchmarks & evaluationMore projects & source code | |
| Minutes Jev voice evals — repogithub.com | — | Synthetic qualification script in the Minutes meeting-memory app that tests Jev on seven Choice decisions its voice path needs, such as attendee constraints, semantic recall, verified pastes, stale targets and prompt injection. | Benchmarks & evaluationMore projects & source code | |
| Miravegithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| Mirave — READMEgithub.com | 0 | Mirave is an open research attempt to reproduce Jev/System One with separate model weights and a mock server; it is not TypeSafe’s official SDK or API. | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 4 | Astro-Han/jev-harness — A coding agent that filters every tool result through Jev before the model sees it, with an A/B harness measuring pass@1 and cost against the unfiltered control · model_id · Python | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 41 | JoshuaSP/open-jev — Typed JSON inference with DiffusionGemma, with Every and Jev benchmark results · model_id · Python | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 79 | NanmiCoder/jev-arena — Jev 模型介绍与实测:通过 Choice / Score / Noul 将自然语言转为带类型的判断与概率,用于分类、评分和路由;支持与 DeepSeek 等模型对比评论打标、速度与结果,含 CSV/Excel 导入、原速回放与离线报告。 · model_id · JavaScript | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 1 | alperenerol/jev-1.13-mini-benchmark — Mini benchmark of TypeSafe's jev-1.13 structured decision model (OpenRouter Decisions API) on labeled support-triage: noul/choice/score, consistency, cost, lessons learned · model_id · Python | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 1 | fabricioism/jev-expirements — A repository for jev-expirements · model_id · TypeScript | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 1 | finetuningsingh/jev-chatbot — Experiment: using TypeSafe Jev as a chatbot by choosing replies one letter or word at a time · model_id · JavaScript | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 8 | kotoba-lang/typed-decisions — Jev-shaped typed-decision model (state + Choice/Score/Noul questions -\> calibrated probabilities, one pass) on ModernBERT / DeBERTa / LLaDA-MoE, with measured latency, accuracy, calibration and training cost · model_id · Python | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 1 | zhuyansen/jev-cold-start-prior — Can a TypeSafe Jev prior read from a README on day one predict which new agent-skill repos gain stars? Zero-shot Jev ≈ a text model trained on ~150 labels; best used as a feature. Prospective test running. · model_id · Python | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 1 | zhuyansen/jev-issue-pulse — Can a Jev-labelled GitHub issue stream catch a broken release before the fix? No at daily cadence (null, n=7). Per issue, Jev matches triage labels far better than keywords or sentiment. · model_id · Python | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 2 | zhuyansen/jev-news-cold-start — Cross-domain check on MIND news: a zero-shot Jev headline prior is worth ~500 labelled articles, adds +0.069 ρ as features, and lifts a Thompson-sampling cold start by 25%. · model_id · Python | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 9 | zhuyansen/jev-search-rerank-eval — Does a TypeSafe Jev rerank beat embedding search? Graded relevance eval (9,831 pairs, 164 zh/en queries) over the Agent Skills Hub catalog, with the judge-circularity bias measured. · model_id · Python | Benchmarks & evaluationMore projects & source code | |
| model_idgithub.com | 1 | zhuyansen/jev-support-pulse — Does a Jev-labelled support-tweet stream spike before a brand admits an outage? At equal false alarms it catches 17 vs 10 incidents (volume), ~4h ahead; a good keyword list is almost as good. · model_id · Python | Benchmarks & evaluationMore projects & source code | |
| no-hallucination — READMEgithub.com | 3 | A set of RAG experiments compares quote-checking, Jev guard and routing strategies; its README reports that Jev did not block the remaining unsupported answers in one test. | Benchmarks & evaluationMore projects & source code | |
| OmarMujahid/jev-decision-benchgithub.com | 2 | An independent benchmark of TypeSafe's Jev, a model that does not write text. You send it some content and a list of typed questions (yes/no, pick one option, rate on a scale) and | Benchmarks & evaluationMore projects & source code | |
| OneJevgithub.com | — | Open models: multimodal System One model in four sizes (0.8B to 27B); typed questions about a screenshot, photo, video or text get a calibrated probability for every option in one forward pass. | Benchmarks & evaluationMore projects & source code | |
| Open JEVgithub.com | 35 | Research preview inspired by Jev that makes local bilingual English and Chinese probability decisions over context, questions and candidate answers, currently on a frozen Qwen3-4B-Instruct backend with a web UI. | Benchmarks & evaluationMore projects & source code | |
| Open Medical Jevgithub.com | — | Medical evaluation: two frozen local readers answer one Noul-style yes/no probability per exam option, a fit-free router auto-releases items above the combined-confidence gate and escalates the rest, and a split-conformal candidate set bounds the error - landing within 2 points of hosted Jev on three 600-item national licensing exams with no fine-tuning, no distillation and no corpus. | Benchmarks & evaluationMore projects & source code | |
| open-jev — READMEgithub.com | 41 | A DiffusionGemma harness explores typed JSON decisions and explicitly identifies itself as an independent, non-reproduction of Jev. | Benchmarks & evaluationMore projects & source code | |
| open-jev-deberta-v3-large — repogithub.com | 8 | Open Jev-shaped model on DeBERTa-v3-large that answers any number of choice, score and noul questions about one state in a single pass; 0.854 in-domain and 0.690 out-of-domain accuracy, 28 ms for 10 questions. | Benchmarks & evaluationMore projects & source code | |
| open-jev-typed-decision-enginegithub.com | 44 | Open reproduction of TypeSafe Jev: a 150M typed decision engine (noul/choice/score in one non-autoregressive pass, calibrated confidence). 0.697 vs Jev's 0.727, 2.5x better calibrated, 4x faster, free. Trains on a Colab T4 in 30 min. | Benchmarks & evaluationMore projects & source code | |
| open-jev-typed-decision-engine — READMEgithub.com | 44 | An independent 150M encoder for typed state questions reports its own benchmark comparison with TypeSafe Jev. | Benchmarks & evaluationMore projects & source code | |
| open-system-one-bench — projectgithub.com | 0 | Per-item predictions for 10,000 classification and routing decisions from six stacks, including typesafe/jev and Laya on identical items, so significance tests can be rerun without API spend. | Benchmarks & evaluationMore projects & source code | |
| openjevgithub.com | 1,307 | Runs a Jev-like interface on a local model by reading option logits instead of generating text. | Benchmarks & evaluationMore projects & source code | |
| openjev — READMEgithub.com | 35 | Open JEV identifies itself as an independent Jev-inspired research alternative and documents a frozen Qwen3 backend with a local candidate-decision interface. | Benchmarks & evaluationMore projects & source code | |
| openjev-experimentsgithub.com | 1 | Self-contained script that runs AlexWortega/openjev, a Qwen3.5-4B cross-encoder, locally for natural language inference, multiple-choice reranking and a CLI. | Benchmarks & evaluationMore projects & source code | |
| openjev-lmgithub.com | 1 | open-Jev LM arm: Qwen2.5-0.5B + LoRA reproducing a hosted decision model's judgment at 92.9% on hand-labelled gold - trained overnight on a 6-vCPU CPU-only host, $0/call. Paper, corpora, harnesses, receipts. | Benchmarks & evaluationMore projects & source code | |
| padflow-jev-evalsgithub.com | 1 | Public benchmark of typed decisions a land-development SaaS makes in production, such as routing incoming documents to projects, with JSON schemas, anonymized labeled rows and a runner for OpenAI-compatible models. | Benchmarks & evaluationMore projects & source code | |
| padflow-jev-evals — READMEgithub.com | 1 | This benchmark publishes PadFlow decision examples so TypeSafe Jev can be measured alongside other candidate decision models. | Benchmarks & evaluationMore projects & source code | |
| paper-package — READMEgithub.com | 1 | A research archive contains a Jev decision study and its reproducibility materials; the README reports a narrow result on one split, not general model superiority or a result reproduced here. | Benchmarks & evaluationMore projects & source code | |
| pi-jev-toolsgithub.com | 0 | TypeSafe ranking, classification, retrieval and structured-decision tools for Pi coding agents | Benchmarks & evaluationMore projects & source code | |
| PocketJevgithub.com | 1 | Experimental iOS app that runs Qwen3-VL on-device via MLX and turns a camera frame, a question and 2-26 options into relative scores read from next-token logits in about a second, with no photos saved. | Benchmarks & evaluationMore projects & source code | |
| pokertools-arena/pokertools-arena.github.iogithub.com | 1 | A browser-first AI poker benchmark. Seat Jev and OpenAI-compatible models at the same no-limit Texas Hold'em table, watch every card and decision as a spectator, and let the tournament run autonomously until one model wins. | Benchmarks & evaluationMore projects & source code | |
| primitivesgithub.com | 303 | hr98w/jev-visual — An educational Jev-like visual inference experiment on Apple Silicon: shared context, direct candidate scoring, and local visual demos. · primitives · Python | Benchmarks & evaluationMore projects & source code | |
| qwen-rlcdgithub.com | 4 | Prototype Jev-style decision model on Qwen3.5-0.8B-Base that prefills the state once, forks the cache per question-answer branch and reads calibrated distributions; inference works, training is on hold. | Benchmarks & evaluationMore projects & source code | |
| qwen-rlcd — READMEgithub.com | 4 | README describes this Qwen3.5 project as a System One decision-model prototype inspired by TypeSafe Jev. | Benchmarks & evaluationMore projects & source code | |
| reflex-jevgithub.com | 0 | Training demonstration of Jev in a dispatch services command center scenario. | Benchmarks & evaluationMore projects & source code | |
| reflex-jev — READMEgithub.com | 0 | An interactive dispatch-center demo assigns typed judgments to Jev and policy decisions to ordinary code; without a key or in replay mode, it plays committed recordings. | Benchmarks & evaluationMore projects & source code | |
| Research Deskgithub.com | 2 | Market-analysis demo that reads live yfinance company profiles and headlines into ranked, grounded and routed trade ideas, with a Requests tab showing the state and questions behind every number. | Benchmarks & evaluationMore projects & source code | |
| RISC-jeVgithub.com | 1 | Hack that runs compiled C on a simulated SERV RISC-V CPU whose logic gates are lookup tables built from Jev's answers to 22 Boolean gate/input combinations, with a trace of clock, signals, requests and cost. | Benchmarks & evaluationMore projects & source code | |
| RISC-jeV — READMEgithub.com | 1 | An experimental simulated-CPU project turns Jev’s Boolean answers into cached lookup tables for logic gates; the README notes there are no correctness checks. | Benchmarks & evaluationMore projects & source code | |
| rogeriochaves/jev-experimentsgithub.com | 1 | Results page: https://claude.ai/code/artifact/f00ee126-9554-4e2f-b2e7-1fc86c066aa9 | Benchmarks & evaluationMore projects & source code | |
| roverlab — READMEgithub.com | 0 | This browser rover sandbox explores autonomous decisions with a selectable TypeSafe controller; the README describes an experiment, not validated robot performance. | Benchmarks & evaluationMore projects & source code | |
| RSI-Jevgithub.com | 42 | Typed-decision models (noul / choice / score) trained by a self-improving loop of AI agents — checkpoints, the code that produced them, and every version that failed. | Benchmarks & evaluationMore projects & source code | |
| RSI-Jev — READMEgithub.com | 42 | A research system whose AI-agent experiment loop builds Jev-style System One models returning probabilities over typed options rather than prose. Its current documented release also accepts images. Listed for the alternative-model research and released artifacts, without certifying self-improvement, latency, benchmark scores or full Jev API compatibility. | Benchmarks & evaluationMore projects & source code | |
| rubikjev — READMEgithub.com | 5 | RubikJev describes Jev rating scramble chaos with meme tiers and a star difficulty rating. | Benchmarks & evaluationMore projects & source code | |
| ruffood/jev-reality-checkgithub.com | 1 | Jev (TypeSafe AI) 可复现实测:算术、计数、日期、零幻觉、把握度校准,中英对照 | Benchmarks & evaluationMore projects & source code | |
| sandroandric/JevGramgithub.com | 1 | AI detection in research papers with Jev | Benchmarks & evaluationMore projects & source code | |
| scarif-labs/jev-software-decision-benchmarkgithub.com | 1 | Reproducible benchmark evaluating JEV as a software decision primitive for dependency-update automation under distribution shift. | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 2 | 0xnairb/research_desk — TypeSafe Jev demonstration for new analyzation — experimenting with Jev for fast analysis of news and tickers · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 21 | AbdelStark/jev-benchmarks — Probability-aware evaluation for typed decision models: calibration, selective risk, latency, and reproducible benchmarks. · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 2 | EmreKaplaner/rag-jev — Make room for useful evidence. Inspectable context selection for RAG, with Jev reranking and open benchmark studies. · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 2 | FirasSX914/Janus — Measure when to use Jev and other models on your data, then route accordingly. · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 1 | JGalego/Jevs-Garage — A garage full of tiny experiments for building critical systems with System One & Jev 🔧🧠⚡ · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 1 | TanayPadar/gpt-vs-jev — Compare GPT generated language with JEV structured Noul decisions on the same input. · sdk · TypeScript | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 26 | Zaious/jev-capability-atlas — Independent, evidence-based map of when TypeSafe's Jev actually holds up vs. breaks down — real API-call receipts, not a leaderboard. 中文為主的雙語 repo。 · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 4 | adhyaay-karnwal/jev-chat — A chatbot from typed Jev decisions: hierarchical speculative decoding over System One probabilities. · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 1 | danielhirt/jev-lab — Experiments on TypeSafe Jev (System One decision model) via OpenRouter: repeatability, perturbation, and LLM baseline comparison · sdk · TypeScript | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 14 | erendikmenn/jev-rag-benchmark — Reproducible benchmark for measuring Jev reranking quality, latency, and cost in RAG · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 3 | jimmyliao/jev-storyboard-lab — Google ADK vs Microsoft Agent Framework for structured-output agents, with TypeSafe Jev as a vendor-neutral QC gate · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 11 | justinhe16/trade-jev — Backtest Jev (TypeSafe) as a BUY/SELL/HOLD trader on NQ L10 order-book data · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 3 | kyrylosyzonenko/jev-browse — jev-browse is an unofficial project and isn't affiliated with TypeSafe or Vercel. · sdk · JavaScript | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 6 | mahlernim/jev-korean-benchmark — Reproducible early-access evaluation of Jev on Korean understanding and medical text, with runtime and cost evidence · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 1 | naveenreddy61/jev-experiments — experiments with system one model jev · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 7 | opaielsheikh/ai-elo-ranker — High-speed recursive AI Elo tournament engine powered by Jev and Swiss matchmaking · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 1 | sandroandric/JevGram — AI detection in research papers with Jev · sdk · Python | Benchmarks & evaluationMore projects & source code | |
| sdkgithub.com | 1 | scarif-labs/jev-software-decision-benchmark — Reproducible benchmark evaluating JEV as a software decision primitive for dependency-update automation under distribution shift. · sdk · TypeScript | Benchmarks & evaluationMore projects & source code | |
| security-sandbox-jevgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| security-sandbox-jev — READMEgithub.com | 0 | This sandbox places a gateway before actions in a fabricated company and uses Jev to judge whether each agent action should be allowed; it does not demonstrate production security efficacy. | Benchmarks & evaluationMore projects & source code | |
| semantic-microscope — READMEgithub.com | 0 | Semantic Microscope labels document sentences with probability-bearing judgments from Jev and provides a document-wide view. | Benchmarks & evaluationMore projects & source code | |
| SHADE-Arena Jev monitorgithub.com | 0 | SHADE-Arena fork evaluating Jev as a monitor for covert agent sabotage against Gemini 2.5 Flash and Pro; as a per-action gate Jev reached AUC 0.97 on 353 tool calls, but only 0.78 on whole transcripts. | Benchmarks & evaluationMore projects & source code | |
| shade-arena-jev-monitor — READMEgithub.com | 0 | This README reports a preliminary Jev sabotage-evaluation study using 13 hand-built transcripts across three tasks; treat the results as design signals, not benchmark numbers. | Benchmarks & evaluationMore projects & source code | |
| simonmesmith/jev-bbq-experimentgithub.com | 1 | Reproducible evaluation of TypeSafe Jev on all 58,492 BBQ questions: accuracy, stereotype bias, uncertainty, cost and latency. | Benchmarks & evaluationMore projects & source code | |
| Six experiments on Jev's real limits — repo2github.com | 9 | Chinese write-up of six experiments across five open-source repos: Jev hits 0.83 AUC on tables with meaningful columns but 0.46 on hashed CTR data, and works best as a feature added to a baseline. | Benchmarks & evaluationMore projects & source code | |
| SmartMoney-Cub Jev layer — repogithub.com | 24 | Optional Jev judgment layer in a read-only trading journal and review harness that asks noul, choice, and score questions about filings and statements; on its 240-case finance benchmark Jev scores 78.43%. | Benchmarks & evaluationMore projects & source code | |
| snake-arena-jev-vs-llmsgithub.com | 0 | How many decisions can a model make in a minute, and what do they cost? Jev, a System One decision model, raced against six LLMs on the same Snake boards. | Benchmarks & evaluationMore projects & source code | |
| snake-arena-jev-vs-llms — READMEgithub.com | 0 | A README-described one-minute Snake benchmark compares Jev with six general-purpose LLMs on the same small decisions; results were not independently reproduced here. | Benchmarks & evaluationMore projects & source code | |
| sstehniy/jev-calculatorgithub.com | 1 | iOS 6-inspired Jev calculator demo with a lifetime API budget | Benchmarks & evaluationMore projects & source code | |
| stuntdouble — READMEgithub.com | 1 | A shadow-comparison proxy preserves Jev's response for the application while recording alternative models' decisions for comparison. | Benchmarks & evaluationMore projects & source code | |
| superradcompany/multiverse-of-madnessgithub.com | 2 | Jev and Microsandbox explore alternate game futures with a reusable TypeScript learning harness | Benchmarks & evaluationMore projects & source code | |
| Sys1Cal-v1 — READMEgithub.com | 0 | Benchmark constructs problems with known probabilities to assess Jev-like Noul, Choice and Score fidelity across equivalent input representations. | Benchmarks & evaluationMore projects & source code | |
| sysadarsh/zerosweepgithub.com | 2 | Autonomous System-One Triage Engine & Benchmark powered by TypeSafe AI (Jev). 75ms inference, $0 output tokens, and RLCD epistemic safety gates. | Benchmarks & evaluationMore projects & source code | |
| sysone-bench — READMEgithub.com | 6 | An independent benchmark comparing Laya, the TypeSafe Jev API and constrained Qwen decisions on a sealed shared set of typed questions. Listed for comparison methodology, not endorsement of reported scores. | Benchmarks & evaluationMore projects & source code | |
| system-one-benchgithub.com | 0 | Hands-on experiments with open-source System One (Jev-style) decision models on an Apple M5: a zero-shot tweet-sentiment benchmark against local LLMs, and a study of how they play Snake. | Benchmarks & evaluationMore projects & source code | |
| system-one-bench — READMEgithub.com | 0 | Hands-on benchmark tests open Jev-style models zero-shot on tweet sentiment, Snake decisions and Finnish comprehension, with task reports and comparison baselines. | Benchmarks & evaluationMore projects & source code | |
| system-one-decision-lab — READMEgithub.com | 0 | Experimental decision lab compares hosted Jev and self-hosted engines on retained domain evidence through one contract, with provider-specific confidence gates rather than a general benchmark. | Benchmarks & evaluationMore projects & source code | |
| system-one-open — READMEgithub.com | 38 | The README describes an independent Gemma-based replica of Jev that returns typed decisions from state in one pass; it does not establish TypeSafe API compatibility. | Benchmarks & evaluationMore projects & source code | |
| system-one-playgroundgithub.com | 0 | Run the new wave of decision models (Laya, Decider, Kev, Jev) side by side on your Mac. Structured input in, calibrated probabilities out, with honest accuracy, calibration, latency and memory numbers. | Benchmarks & evaluationMore projects & source code | |
| system-one-playground — READMEgithub.com | 0 | Decision-model comparison workbench evaluates local models alongside hosted TypeSafe Jev, exposing accuracy, calibration, latency and resource measurements. | Benchmarks & evaluationMore projects & source code | |
| system-one-security-triagegithub.com | 1 | Recorded comparison of Jev, Terra, and Opus on 100 synthetic security-triage cases, five passes each, with a static inspectable dashboard | Benchmarks & evaluationMore projects & source code | |
| system-one-security-triage — READMEgithub.com | 1 | A recorded security-evidence triage comparison including TypeSafe Jev through its SDK. It is an evaluation/demo resource, not evidence that Jev safely resolves all security findings. | Benchmarks & evaluationMore projects & source code | |
| terrarium — READMEgithub.com | 0 | A creature sandbox lets live TypeSafe System One choose a button while deterministic code runs the world; without a key, the page uses a mock brain. | Benchmarks & evaluationMore projects & source code | |
| TetraJevgithub.com | — | General decisions: two frozen open-weight readers give four readings per item (letter + per-candidate yes/no), fused fit-free and routed by agreement into auto-release or human review; evaluated across eight decision suites and a RAG reranking pass with published coverage–accuracy curves; no training, and it does not call the TypeSafe API. | Benchmarks & evaluationMore projects & source code | |
| thaiexam-jev-charts — READMEgithub.com | 0 | The repository publishes charts from an evaluation of TypeSafe Jev on Thai standardized exams alongside other language models. | Benchmarks & evaluationMore projects & source code | |
| tiny-jevgithub.com | 0 | Small local judgment model for Japanese text that returns choice, score, and noul answers from Qwen3 yes/no scores with LoRA training, abstaining until business data and calibration are in place. | Benchmarks & evaluationMore projects & source code | |
| tiny-jev — READMEgithub.com | 0 | A local Japanese typed-decision prototype maps Qwen3 yes/no scores to Choice/Score/Noul-shaped outputs; it does not establish Jev API compatibility. | Benchmarks & evaluationMore projects & source code | |
| tiny-llm2github.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| tiny-llm2 — READMEgithub.com | 0 | Learning repository combines small language models with Jev/Laya examples for banking decisions and comparisons, explicitly labelling offline stand-ins as not Jev and non-production demos. | Benchmarks & evaluationMore projects & source code | |
| trade-jev — READMEgithub.com | 11 | This research repository uses TypeSafe Jev as a bounded BUY/SELL/HOLD decision layer over NQ L10 order-book data. | Benchmarks & evaluationMore projects & source code | |
| trading-bot-jevgithub.com | 2 | Medium-term crypto trend follower on 4-hour candles where Jev reads live news and an algorithm decides trades on a 1,000 USDC paper portfolio, built to test what Jev adds over the algorithm alone. | Benchmarks & evaluationMore projects & source code | |
| trading-bot-jev — READMEgithub.com | 2 | A crypto bot reads market news with Jev while a separate algorithm makes trading decisions; the README says it can use a paper portfolio or a real account on request. | Benchmarks & evaluationMore projects & source code | |
| TypeARgithub.com | 8 | Type-safe one-decision-per-token decoding engine for autoregressive LLMs, inspired by Jev. | Benchmarks & evaluationMore projects & source code | |
| typed-decision-benchgithub.com | 0 | Large typed decision model evaluation benchmark and novel calibration standard (calibration.json) for any typed decision inference system, covering 275 distinct use-cases and 27,598 individual decisions for Jev-like open typed decision models. | Benchmarks & evaluationMore projects & source code | |
| typed-decision-bench — READMEgithub.com | 0 | This small benchmark compares Jev with OpenJev, Laya, and local/hosted models on shared zero-shot English/French classification tasks; the authors limit it to one dataset and question type, and the metrics were not reproduced. | Benchmarks & evaluationMore projects & source code | |
| typed-decision-bench — READMEgithub.com | 0 | typed-decision-bench is an evaluation and calibration suite for systems implementing the System One typed-decision API, including Jev-compatible requests. Its README describes large held-out suites; published results were not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| typed-decisions — READMEgithub.com | 8 | A lightweight experiment reproduces the Jev-style Choice, Score, and Noul shape while explicitly treating Jev as a reference rather than an implementation specification. | Benchmarks & evaluationMore projects & source code | |
| typed_evals — READMEgithub.com | 13 | Typed Evals uses Jev to judge generated responses and recorded agent executions, and can check proposed tool calls. | Benchmarks & evaluationMore projects & source code | |
| typesafe-aigithub.com | 0 | TypeSafe AI: structured AI judgments as programming primitives | Benchmarks & evaluationMore projects & source code | |
| typesafe-ai — READMEgithub.com | 0 | Text-processing pipeline uses Jev to classify recommendation blocks and verify extracted provider names before assembling structured endorsement results. | Benchmarks & evaluationMore projects & source code | |
| typesafe-ai-benchmarkgithub.com | 39 | Side-by-side benchmark of Qwen 3.8 27B structured output on Cerebras versus Jev across seven synthetic workloads, recording mistakes, latency, tokens, and estimated cost. | Benchmarks & evaluationMore projects & source code | |
| typesafe-ai-benchmark — READMEgithub.com | 39 | A synthetic-workload benchmark compares TypeSafe Jev with a constrained Qwen structured-output path and a local tool-calling model; its results are author-reported. | Benchmarks & evaluationMore projects & source code | |
| typesafe-ai-firewallgithub.com | 1 | Shadow-mode validation harness for a pre-execution firewall on AI agent tool calls (TypeSafe/Jev). Real run, findings in report.md. | Benchmarks & evaluationMore projects & source code | |
| typesafe-ai-playground — READMEgithub.com | 4 | TypeSafe AI Playground is for Jev experiments; the README says OpenRouter takes precedence when both provider keys are set. | Benchmarks & evaluationMore projects & source code | |
| typesafe-ai_smoke-testgithub.com | 0 | Small router that asks Jev two yes/no questions per prompt to send it to a work store, a life store or both, with a 30-case hand-written eval where all 30 routed correctly. | Benchmarks & evaluationMore projects & source code | |
| typesafe-arenagithub.com | 3 | A playground for TypeSafeAI's Jev Model. | Benchmarks & evaluationMore projects & source code | |
| typesafe-arena — READMEgithub.com | 3 | typesafe-arena is a local mirror of TypeSafe documentation with a Lume search index, rather than a Jev-powered decision app. | Benchmarks & evaluationMore projects & source code | |
| typesafe-guardrails — READMEgithub.com | 1 | A reproducible dealership-chatbot jailbreak demo compares guardrail modes, with Jev screening decision boundaries in the agent. | Benchmarks & evaluationMore projects & source code | |
| typesafe-jev-calibrate-for-code-review — READMEgithub.com | 3 | This practitioner report describes code-review rulesets in Jev for both comments and code. | Benchmarks & evaluationMore projects & source code | |
| typesafe-jev-dojo — READMEgithub.com | 0 | The Jev dojo distinguishes a simulated no-key mode from its LIVE mode, which makes real Jev calls. | Benchmarks & evaluationMore projects & source code | |
| typesafe-jev-traffic-demo — READMEgithub.com | 1 | This local traffic-junction proof of concept asks Jev to choose a phase priority from open sensor data, but README explicitly says it is a simulation that cannot control signals. | Benchmarks & evaluationMore projects & source code | |
| typesafe-oraclesgithub.com | 0 | Evaluating TypeSafe's System One primitives (Choice/Score/Noul) — where a typed oracle beats an LLM call | Benchmarks & evaluationMore projects & source code | |
| typesafe-playground — READMEgithub.com | 13 | A set of interactive TypeSafe Jev experiments ranges from support-message judgments to decisions in a 3D car scene; the README says these use real API calls, not polished benchmarks. | Benchmarks & evaluationMore projects & source code | |
| typesafe-showcase — READMEgithub.com | 1 | This Next.js showcase presents TypeSafe System One (Jev) demos; each visitor supplies their own API key, so the app does not provide a shared key. | Benchmarks & evaluationMore projects & source code | |
| typesafe-vs-deepseekgithub.com | 1 | Side-by-side comparison of Jev and DeepSeek flash on speed, tokens, cost and accuracy across invoice extraction, email classification and reranking, plus fraud, guardrail and reconciliation pipelines. | Benchmarks & evaluationMore projects & source code | |
| Typesafe_chess_evalgithub.com | 0 | An evaluation of typesafe AI chess. As it turns out, the AI isn't doing really well even though chess is not a particularly open-ended game. Still, it's only a prototype and this probably wasn't optimzied for games. | Benchmarks & evaluationMore projects & source code | |
| Typesafe_chess_eval — Chess.pygithub.com | 0 | The chess prototype calls TypeSafe system_one with Choice questions about outcomes and moves; the README does not pin a Jev model version. | Benchmarks & evaluationMore projects & source code | |
| ui-generator-instinct-jev — READMEgithub.com | 7 | Instinct is a UI demo where Jev answers typed questions; the README says Jev does not generate code or copy. | Benchmarks & evaluationMore projects & source code | |
| uipath-maestroflow-jevgithub.com | 0 | UiPath Maestro Flow x TypeSafe Jev - a card dispute triage bench comparing a calibrated typed-decision model against an inline AI agent on identical work, with latency and cost measured at runtime. | Benchmarks & evaluationMore projects & source code | |
| uipath-maestroflow-jev — READMEgithub.com | 0 | The README presents a UiPath Maestro demo that routes ten card-dispute emails through both Jev and an inline UiPath-agent path; reported runtime metrics were not reproduced here. | Benchmarks & evaluationMore projects & source code | |
| What Jev Is Missinggithub.com | 0 | Position paper arguing that per-hop calibration does not compose across decision pipelines and that typed answers destroy vagueness, proposing Hidden-Markov and fuzzy primitives and an architecture called BSF-S1. | Benchmarks & evaluationMore projects & source code | |
| what-is-jevgithub.com | 1 | Independent, source-linked research on TypeSafe AI's Jev (System One), with 947 rubric-scored public repositories, recurring patterns, datasets, and bilingual documentation. | Benchmarks & evaluationMore projects & source code | |
| what-is-jev — READMEgithub.com | 1 | what-is-jev is independent research that links sources about TypeSafe Jev and surveys projects using it; it states repository scores are not project benchmarks. | Benchmarks & evaluationMore projects & source code | |
| WHAT-s-Up-jevgithub.com | 0 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| WHAT-s-Up-jev — READMEgithub.com | 0 | This repository is an evaluation log with notes and code about Jev, not a product or recommendation. | Benchmarks & evaluationMore projects & source code | |
| what-the-jev — READMEgithub.com | 4 | A showcase collection introducing Jev/TypeSafe System One tasks. Included as a learning resource, not independent verification of its examples’ claimed effectiveness. | Benchmarks & evaluationMore projects & source code | |
| WindTunnel Jev experimentsgithub.com | — | WebMCP benchmark runs where Jev picks browser actions and Mercury writes arguments and answers: via WebMCP the pair solved 49/49 tasks at a $0.0011 median per run, versus 25/49 through DOM control. | Benchmarks & evaluationMore projects & source code | |
| WindTunnel Jev experiments — repogithub.com | — | WebMCP benchmark runs where Jev picks browser actions and Mercury writes arguments and answers: via WebMCP the pair solved 49/49 tasks at a $0.0011 median per run, versus 25/49 through DOM control. | Benchmarks & evaluationMore projects & source code | |
| zhuyansen/jev-cold-start-priorgithub.com | 1 | Can a TypeSafe Jev prior read from a README on day one predict which new agent-skill repos gain stars? Zero-shot Jev ≈ a text model trained on ~150 labels; best used as a feature. Prospective test running. | Benchmarks & evaluationMore projects & source code | |
| zhuyansen/jev-issue-pulsegithub.com | 1 | Can a Jev-labelled GitHub issue stream catch a broken release before the fix? No at daily cadence (null, n=7). Per issue, Jev matches triage labels far better than keywords or sentiment. | Benchmarks & evaluationMore projects & source code | |
| zhuyansen/jev-news-cold-startgithub.com | 2 | Cross-domain check on MIND news: a zero-shot Jev headline prior is worth ~500 labelled articles, adds +0.069 ρ as features, and lifts a Thompson-sampling cold start by 25%. | Benchmarks & evaluationMore projects & source code | |
| zio-evalsgithub.com | 3 | TypeSafe Jev ecosystem repository. | Benchmarks & evaluationMore projects & source code | |
| zio-evals — READMEgithub.com | 3 | zio-evals provides JevJudge as a TypeSafe Jev typed-binary grader, configured as an alternative to a generative judge. | Benchmarks & evaluationMore projects & source code | |
| zzzzzec/jevsortgithub.com | 1 | Jev-powered integer sorting experiment: serial selection versus parallel rank prediction. | Benchmarks & evaluationMore projects & source code | |
| ← Back to Awesome Jevgithub.com | 24 | Listed by the source without a separate description. | Benchmarks & evaluationMore projects & source code |
More posts & discussions43 of 43 matches
| Resource | Stars | Description | Category | Save |
|---|---|---|---|---|
| $200 test of Jevx.com | — | Chinese evaluation that spent $200 running Jev against GPT-4.1-mini and GPT-5.6 Sol, Terra and Luna on 2,390 questions, including Chinese-language tasks, concluding the 100x cheaper claim roughly holds but 100x faster does not. | Benchmarks & evaluationMore posts & discussions | |
| 1,000 papers sorted then auditedx.com | — | Open rebuild of a Jev paper map: Jev sorted 1,000 AI papers into 24 topics for $0.0585, and Opus 5 as judge agreed on 85 of 100 labels at 153x the cost and 1.9s versus 57ms per paper. | Benchmarks & evaluationMore posts & discussions | |
| A different kind of model for AI observability — discussionnews.ycombinator.com | — | Labels 10,000 agent traces for status and sentiment in about 17 minutes for under a dollar and compares accuracy with an evaluation model. | Benchmarks & evaluationMore posts & discussions | |
| Can Jev Be a Better Agent Evaluator? — demox.com | — | LangChain tests Jev as a judge against LLM judges for agent evaluation in LangSmith, comparing accuracy, repeatability, latency and cost. | Benchmarks & evaluationMore posts & discussions | |
| Choice vs Noul trolley problemsx.com | — | Runs a series of trolley problems through Jev as both a Choice and a Noul to test whether the question type changes its judgment. | Benchmarks & evaluationMore posts & discussions | |
| Code review benchmarkx.com | — | Benchmark of Jev scoring raw Git diffs against a GLM + Grok + Gemini ensemble reviewer: zero false positives, ~50x faster, ~100x cheaper, with 75% bug recall. | Benchmarks & evaluationMore posts & discussions | |
| DecaState Jev vs frontier LLMswww.reddit.com | — | Benchmark of Jev against four frontier LLMs on 36 support-triage decisions from US-West and Singapore: about 176x cheaper and 9x faster than GPT-6 Astra, with 0 type errors. | Benchmarks & evaluationMore posts & discussions | |
| Does Jev have politics? (Yes) — discussionnews.ycombinator.com | — | Runs Political Compass propositions through Jev as Choice questions and finds stable political positions that match most LLMs, arguing it is a new probabilistic mask on the same underlying model. | Benchmarks & evaluationMore posts & discussions | |
| Hermes Agent compaction scorecard — demox.com | — | Tests a Jev-based context-compaction plugin against Hermes' own compressor and recommends against it: far cheaper and faster per compaction, but it kept twice the context and ranked tool results no better than recency. | Benchmarks & evaluationMore posts & discussions | |
| HiringCafe resume-job relevance benchmarkx.com | — | Thread benchmarking Jev on resume-to-job-description relevance scoring for HiringCafe, a job search app serving 2.5 million monthly active users. | Benchmarks & evaluationMore posts & discussions | |
| Jev adoption on Vercel AI Gateway — discussionnews.ycombinator.com | — | First-day adoption data from Vercel AI Gateway: within 24 hours nearly 13% of paid teams were using Jev, about 2x the GPT-5.6 family and more than 6x Fable 5.1's share. | Benchmarks & evaluationMore posts & discussions | |
| Jev as a reranker on DL19/DL20x.com | — | IR researchers test Jev as a pointwise, pairwise, setwise and listwise reranker over the top 100 BM25 results on TREC DL19 and DL20 and find it good and cheap. | Benchmarks & evaluationMore posts & discussions | |
| Jev guesses what I drew — discussionnews.ycombinator.com | — | Experiment that turns 400 hand-drawn sketches into text and asks Jev to name them: it beat chance by a wide margin, lost to Claude Sonnet 5, and answered airplane for more than half. | Benchmarks & evaluationMore posts & discussions | |
| Jev in a security pipelinex.com | — | Results table from a production security pipeline: Jev reached 99.3% accuracy at 0.259s latency and $0.026 per 1K, versus Gemini and open models that were slower and costlier. | Benchmarks & evaluationMore posts & discussions | |
| Jev is the fish at the poker table — discussionnews.ycombinator.com | — | Poker probe showing Jev swinging 15 to 30 points when the same hand is relabelled and betting against a known made flush in 16 of 16 runs, as a warning against deploying it unevaluated. | Benchmarks & evaluationMore posts & discussions | |
| Jev on 12 real use casesx.com | — | Hands-on review across 12 automations: 1,000 emails through seven decision rules for about nine cents in six seconds after parallelizing, plus comments, meetings, clips and a BTC paper trader. | Benchmarks & evaluationMore posts & discussions | |
| Jev on LLM-Wikiracex.com | — | Researchers ran Jev on their LLM-Wikirace benchmark (450 games, 8 hours, under $1 total) and found it fast and cheap but short on the world knowledge and planning of frontier LLMs. | Benchmarks & evaluationMore posts & discussions | |
| Jev on ScopeJudgex.com | — | Security firm dreadnode ran Jev against its ScopeJudge benchmark for agent scope violations and found it competitive with leading LLM judges, at pennies per thousand checks and 130 ms average responses. | Benchmarks & evaluationMore posts & discussions | |
| Jev robotics benchmarkx.com | — | Benchmark giving Jev a robot body across 120 real and simulated navigation and spatial-reasoning tasks, graded on speed, cost, tokens, collisions and path quality against Dimcode, Astra, Fable, Opus and 5.6. | Benchmarks & evaluationMore posts & discussions | |
| Jev vs DeepSeek ticket routingx.com | — | Side-by-side routing of 500 real e-commerce support tickets: Jev cleared them in 83 seconds for $0.01, while DeepSeek V4.1 Flash had done 173 for $0.06 when stopped. | Benchmarks & evaluationMore posts & discussions | |
| Jev vs. classical ML — discussionwww.reddit.com | — | Eight datasets against eleven classical pipelines, strong on text such as IMDb reviews and weak on tabular data, with notebooks. | Benchmarks & evaluationMore posts & discussions | |
| Jev web form filling benchmarkwww.reddit.com | — | Harness testing Jev on three real web forms, from a 13-step insurance wizard to a Lever job application; at $0.001-$0.006 per form it finished 0 of 13 steps and 7 of 11 fields. | Benchmarks & evaluationMore posts & discussions | |
| jev-eval-agent — demox.com | — | Experiment with a personal-assistant agent and 100 mocked tools that counts the steps needed when the LLM picks tools itself versus when Jev picks the tool and the LLM only fills arguments. | Benchmarks & evaluationMore posts & discussions | |
| jev-rag-benchmark — demox.com | — | Reproducible benchmark of Jev as the reranker in a small RAG system on 1,044 Turkish XQuAD questions, measuring quality, latency and cost with the same 20 candidates given to every reranker. | Benchmarks & evaluationMore posts & discussions | |
| jev-secret-detectionx.com | — | Measures how well TypeSafe's RLCD-Jev model spots real secret credentials in file snippets | Benchmarks & evaluationMore posts & discussions | |
| jev-secret-detection — discussionwww.reddit.com | — | Benchmark of how well Jev spots real, usable secret credentials in 100 file snippets plus edge and config sets, reporting accuracy, AUC, and recall with no regex or provider verification. | Benchmarks & evaluationMore posts & discussions | |
| jevlikex.com | — | Train a small model that chooses among a changing list of text options, one probability per option in a single pass. Includes Doom, chess, and Wikispeedia demos. | Benchmarks & evaluationMore posts & discussions | |
| jevmlxx.com | — | Jev-style parallel constrained decisions for any MLX model on Apple Silicon. Typed, schema-valid JSON in one forward pass. | Benchmarks & evaluationMore posts & discussions | |
| Judgment arenax.com | — | Video that explains System 1 vs System 2 and races Jev against Claude Opus, Haiku 4.5 and GPT-5.4 Mini on 15 human-labelled questions; Opus got one more right but took 10x longer and cost 146x more. | Benchmarks & evaluationMore posts & discussions | |
| Models watching models — discussionnews.ycombinator.com | — | Flags risky calls across 220,000 agent tool calls, shows that encoding tricks slip past it, and finds worded scales beat 1-to-100 ratings. | Benchmarks & evaluationMore posts & discussions | |
| Models watching models — postx.com | — | Flags risky calls across 220,000 agent tool calls, shows that encoding tricks slip past it, and finds worded scales beat 1-to-100 ratings. | Benchmarks & evaluationMore posts & discussions | |
| OpenRouter Ori Eval judging testx.com | — | OpenRouter's Ori Eval comparison of Jev with popular LLMs as a judge: Jev was more than 5x faster than the next fastest model, and its slowest requests beat every other model's median. | Benchmarks & evaluationMore posts & discussions | |
| Pac-Man follow-upx.com | — | Author’s report that an earlier win could not be reproduced. | Benchmarks & evaluationMore posts & discussions | |
| Rebuilding Nym's agent around Jev — discussionnews.ycombinator.com | — | Replaces seven LLM reviewers and tool selection with Jev, falling back to a larger model when unsure, for a 4.1 to 5.7x speedup. | Benchmarks & evaluationMore posts & discussions | |
| Rewriting rules changes Jev's answersx.com | — | Japanese study of Jev as a judgment layer for a design harness, showing how clearer brand-guideline wording flips its answers; 9 test types, 650 calls and 4,819 judgments. | Benchmarks & evaluationMore posts & discussions | |
| Six experiments on Jev's real limitsx.com | — | Chinese write-up of six experiments across five open-source repos: Jev hits 0.83 AUC on tables with meaningful columns but 0.46 on hashed CTR data, and works best as a feature added to a baseline. | Benchmarks & evaluationMore posts & discussions | |
| Six things I tried with Jev — xx.com | — | Six experiments swapping Gemini for Jev, from script fact-checking (24/24 correct, 0.41 s median) and news-feed ranking to PDF passage reranking (right passage first 7 times vs once) and agent-failure review. | Benchmarks & evaluationMore posts & discussions | |
| System One models in high-throughput data pipelines — discussionnews.ycombinator.com | — | Entity resolution with Jev doing the bulk work and an LLM reviewing, at 226x lower cost; rewriting criteria as "what counts as sufficient evidence" fixed the dev set. | Benchmarks & evaluationMore posts & discussions | |
| Trolley problems: humans vs robotsx.com | — | Video of Jev working through trolley problems, where it chose to sacrifice a human to save robots. | Benchmarks & evaluationMore posts & discussions | |
| TypeARx.com | — | Type-safe one-decision-per-token decoding engine for autoregressive LLMs, inspired by Jev. | Benchmarks & evaluationMore posts & discussions | |
| Typed Decisions benchmarkwww.reddit.com | — | New 400-case benchmark for typed decisions where Jev scores 0.727, near the 0.735 teacher ceiling, against 0.646 for a frozen 149M ModernBERT encoder with small heads. | Benchmarks & evaluationMore posts & discussions | |
| typesafe-ai-benchmark — demox.com | — | Side-by-side benchmark of Qwen 3.8 27B structured output on Cerebras versus Jev across seven synthetic workloads, recording mistakes, latency, tokens, and estimated cost. | Benchmarks & evaluationMore posts & discussions | |
| WindTunnel — demox.com | — | WebMCP browser-agent benchmark of 49 tasks on 8 real sites across 21 configurations, where Jev + Mercury 2.5 tops the composite score, solving 49/49 tasks at $0.0011 median cost per task. | Benchmarks & evaluationMore posts & discussions |
More videos & channels1 of 1 matches
| Resource | Stars | Description | Category | Save |
|---|---|---|---|---|
| jev-labs — videowww.youtube.com | — | TLA+-verified consensus kernel around Jev, generated into Rust and run through 1,680 simulated pharmacy decisions under seeded chaos against the live API, with zero wrong verdicts and more escalations as evidence degrades. | Benchmarks & evaluationMore videos & channels |
More guides & websites60 of 60 matches
| Resource | Stars | Description | Category | Save |
|---|---|---|---|---|
| 48 Japanese calls to Jevzenn.dev | — | Japanese test of 16 synthetic support messages sent 3 times each via OpenRouter with three questions per call: 48 calls at a median 286 ms for about $0.00146 total, with low confidence where labels disagreed. | Benchmarks & evaluationMore guides & websites | |
| A different kind of model for AI observabilityfatliverfreddy.substack.com | — | Labels 10,000 agent traces for status and sentiment in about 17 minutes for under a dollar and compares accuracy with an evaluation model. | Benchmarks & evaluationMore guides & websites | |
| Agent Trace Observability workflow evalevals.typesafe.ai | — | Official workflow eval for triaging finished support-agent traces into follow-ups; Jev scores 71.6% at $0.0003 and 0.5 s per case versus Opus 5's 75.2% at $0.1033 and 27.4 s. | Benchmarks & evaluationMore guides & websites | |
| An early-access test of Jevlindfors.no | — | Scores Norwegian documents with Score questions at about $0.22 per thousand and finds that longer, more detailed questions hurt calibration. | Benchmarks & evaluationMore guides & websites | |
| Armature TypeSafe decision A/B — apparmature.now | — | A/B benchmark inside the Armature agent harness on 20 labeled code-review statements: Jev scored 0.87 overall versus 0.98 for a qwen3.6-27b judge, at 260 ms versus 20,591 ms per call. | Benchmarks & evaluationMore guides & websites | |
| Can a fast AI gate catch chemistry mistakes?frederickparsons.substack.com | — | A literature-claim gate that stopped all 42 bad claims at a 95% threshold, with an honest account of where molecule checks failed. | Benchmarks & evaluationMore guides & websites | |
| Can Jev Be a Better Agent Evaluator?www.langchain.com | — | LangChain tests Jev as a judge against LLM judges for agent evaluation in LangSmith, comparing accuracy, repeatability, latency and cost. | Benchmarks & evaluationMore guides & websites | |
| choxos/jev-reviewer websitejevreviewer.xera.ac | — | Data extraction for systematic reviews, quoted from the papers. Ask a trial report and its supplements your extraction form or a RoB 2, ROBINS-I, QUADAS-2 or TIDieR template; Jev points at the lines, every answer is a verbatim quote with its page, you check it and export the table. Files stay in your browser. | Benchmarks & evaluationMore guides & websites | |
| Convex decision model evals — appconvex-evals.netlify.app | — | Benchmark of 106 decision questions drawn from 90 Convex coding evals that compares Jev via OpenRouter's decisions API against language models, recording probabilities, confidence and cost. | Benchmarks & evaluationMore guides & websites | |
| cua-s1-formshuggingface.co | — | Tiny Jev-like option scorer (706,048 parameters, 2.8 MB) that rates fill, check, click or skip for every form field in one parallel pass, built as the decision layer for cua-driver. | Benchmarks & evaluationMore guides & websites | |
| Customer Service workflow evalevals.typesafe.ai | — | Official workflow eval for choosing a support assistant's next actions on a customer turn; Jev scores 76.0% at $0.0001 and 0.4 s per case versus Opus 5's 72.4% at $0.0579 and 16.6 s. | Benchmarks & evaluationMore guides & websites | |
| Does Jev have politics? (Yes)idlerambling.substack.com | — | Runs Political Compass propositions through Jev as Choice questions and finds stable political positions that match most LLMs, arguing it is a new probabilistic mask on the same underlying model. | Benchmarks & evaluationMore guides & websites | |
| Ensemblr Jev decision layer study — appwww.ensemblr.dev | — | Design proposal and spike for using Jev in a desktop orchestrator for Pi and Claude Code to pick agent roles, rate difficulty and flag duplicates; rejected after a corrected rerun that still did not support production use. | Benchmarks & evaluationMore guides & websites | |
| f00ee126-9554-4e2f-b2e7-1fc86c066aa9claude.ai | — | rogeriochaves/jev-experiments — Results page: https://claude.ai/code/artifact/f00ee126-9554-4e2f-b2e7-1fc86c066aa9 · endpoint · Go | Benchmarks & evaluationMore guides & websites | |
| Fine-tuning side quests Jev could have removedkasra.blog | — | Filters 120,633 training examples with three Nouls in 23 minutes for $3.47. | Benchmarks & evaluationMore guides & websites | |
| genai-craft/openvons websitegenai-craft.com | — | openvons (open-Jev): 有限選択肢に確率で答える判断層 — テキスト / 画像 / 日本語音声コマンド | Benchmarks & evaluationMore guides & websites | |
| GoodWatch Jev fingerprint experiment — appgoodwatch.app | — | Case study from the GoodWatch movie-discovery app testing Jev for 74 per-title trait scores: 10 to 25 times faster but a 2.65-point mean gap to reviewed Qwen scores, so the team decided not to adopt it. | Benchmarks & evaluationMore guides & websites | |
| HAKARI-Bench Jev reranker — leaderboardhuggingface.co | — | Jev integration in HAKARI-Bench, a lightweight IR benchmark over 35+ benchmark groups, that ranks documents by Noul relevance probabilities in listwise or pointwise mode with jev-1.13.0 pinned. | Benchmarks & evaluationMore guides & websites | |
| Hermes Agent compaction scorecard — apphermes-agent.nousresearch.com | — | Tests a Jev-based context-compaction plugin against Hermes' own compressor and recommends against it: far cheaper and faster per compaction, but it kept twice the context and ranked tool results no better than recency. | Benchmarks & evaluationMore guides & websites | |
| HiringCafe resume-job relevance benchmark — apphiringcafe.com | — | Thread benchmarking Jev on resume-to-job-description relevance scoring for HiringCafe, a job search app serving 2.5 million monthly active users. | Benchmarks & evaluationMore guides & websites | |
| https://jev-playground-five.vercel.appjev-playground-five.vercel.app | — | Try it yourself: | Benchmarks & evaluationMore guides & websites | |
| https://jev-playground-five.vercel.app/reportjev-playground-five.vercel.app | — | Full report (Chinese, with charts): | Benchmarks & evaluationMore guides & websites | |
| I ran 40 tickets through Jev and four modelsthoughts.jock.pl | — | Runs 40 support tickets through Jev and four text models: Jev answered in 370 ms for two hundredths of a cent, was 10x faster and 329x cheaper than Claude Fable 5.1, and beat it on customer annoyance. | Benchmarks & evaluationMore guides & websites | |
| iammrduncan/typesafe-ai-benchmark websitehackersintheloop.org | — | This is a LLM Gateway that mimics typesafe ai structured output. Like an imposter Jev. | Benchmarks & evaluationMore guides & websites | |
| Invoice Processing workflow evalevals.typesafe.ai | — | Official workflow eval for deciding whether and how an invoice can be paid; Jev scores 61.8% at $0.0011 and 0.5 s per case versus Opus 5's 78.4% at $0.4856 and 92.1 s. | Benchmarks & evaluationMore guides & websites | |
| Jev 1.13 as a Reward Model — reportgoya4140.github.io | — | Reproducible evaluation of Jev 1.13 as a reward model, LLM judge and process verifier across eight benchmark tracks and 40,940 examples, with an interactive report of 54 SOTA comparisons. | Benchmarks & evaluationMore guides & websites | |
| Jev adoption on Vercel AI Gatewayvercel.com | — | First-day adoption data from Vercel AI Gateway: within 24 hours nearly 13% of paid teams were using Jev, about 2x the GPT-5.6 family and more than 6x Fable 5.1's share. | Benchmarks & evaluationMore guides & websites | |
| Jev AI Testedwww.mindstudio.ai | — | Hands-on test of Jev 1.13.0 on eight synthetic cases covering ticket routing, negation, prompt injection, an 'other' option, and latency, with notes on where it holds up and where it does not. | Benchmarks & evaluationMore guides & websites | |
| Jev Benchmark (ads)huggingface.co | — | Research article comparing Jev with GPT-5.6 Sol and four XGBoost baselines on four synthetic ad outcomes: similar aggregate quality at about 64x lower estimated API cost and 5.4x lower median latency. | Benchmarks & evaluationMore guides & websites | |
| Jev guesses what I drewmikulskibartosz.name | — | Experiment that turns 400 hand-drawn sketches into text and asks Jev to name them: it beat chance by a wide margin, lost to Claude Sonnet 5, and answered airplane for more than half. | Benchmarks & evaluationMore guides & websites | |
| Jev in Korean — siteahn-lab.org | — | Frozen, reproducible 100-question sample check of Jev on Korean text: reading comprehension scored 96 in Korean versus 97 in English, while fine-grained meaning judgments scored 76 versus 80. | Benchmarks & evaluationMore guides & websites | |
| Jev in Search: Three Practical Evaluationszc277584121.github.io | — | Independent experiments on search stopping, memory reranking, and multi-hop relation selection, with implementation links and limitations including private data, unequal sample counts, and a simulated speed illustration. | Benchmarks & evaluationMore guides & websites | |
| Jev research deck — slidesmyokoym.github.io | — | Ongoing Japanese research on Jev and System One models maintained as Markdown slides, a presentation script and an article, backed by a ledger of sources, third-party verification and caveats. | Benchmarks & evaluationMore guides & websites | |
| Jev synthetic survey — articlejjd-lab.github.io | — | Study running Jev and GPT-4.1 as the same 300 synthetic respondents over 24,596 Twin-2K-500 cells, finding that asking yes/no items as a Noul mattered more than the model gap, at a thirty-fourth of the cost. | Benchmarks & evaluationMore guides & websites | |
| Jev web form filling benchmark — datagist.github.com | — | Harness testing Jev on three real web forms, from a 13-step insurance wizard to a Lever job application; at $0.001-$0.006 per form it finished 0 of 13 steps and 7 of 11 fields. | Benchmarks & evaluationMore guides & websites | |
| jev-access-day — appjev-access-day.vercel.app | — | Learning scaffold with an eval harness comparing Jev against an LLM stand-in on 24 real operational decisions, with numbers reproducible from committed run files and an interactive results page. | Benchmarks & evaluationMore guides & websites | |
| Jev: 62.6% asked once, 95% split five wayswww.beri.net | — | Roundup of independent tests: Jev ran 12-27x cheaper than Claude Haiku 4.5 on a 2,000-email phishing test, scored 62.6% with one question and 95.0% split into five, and showed miscalibrated probabilities. | Benchmarks & evaluationMore guides & websites | |
| Jevals — appjevals.com | — | Independent benchmark data comparing Jev with six LLMs on PubMedQA, Banking77 and HelpSteer2 against human labels, covering accuracy, calibration, cost and latency. | Benchmarks & evaluationMore guides & websites | |
| JevArena — appjevarena-lab.vercel.app | — | Bring-your-own-key arena that pits Jev against another judge model from OpenRouter on your question and takes your vote before revealing which was which, along with latency and cost. | Benchmarks & evaluationMore guides & websites | |
| JevBench (Benchmark Heaven) — appbenchmarkheaven.com | — | Benchmark for Jev and Jev-like decision models inside Benchmark Heaven, an LLM price and benchmark comparison site, with public and held-out task splits reported separately. | Benchmarks & evaluationMore guides & websites | |
| JEVfire — appkikoncuo.github.io | — | Jev-inspired parallel decisions for CUDA LLMs on vLLM that score single-token labels with the model's own head and assemble JSON in code, with game demos including in-browser Mario at 71 ms per action. | Benchmarks & evaluationMore guides & websites | |
| JEVQA video-quality evaluationarxiv.org | — | Preprint and public result data evaluating Jev for zero-shot video-quality prediction from metadata and codec features. | Benchmarks & evaluationMore guides & websites | |
| largitdata: "Jev System One Model open-source benchmark"www.largitdata.com | — | multi-turn RAG routing — Gemma 4 31B vs Jev vs open alternatives; latency advantage for Jev. [independent] | Benchmarks & evaluationMore guides & websites | |
| LLM Chess: Jev results — appmaxim-saplin.github.io | — | Long-running chess benchmark for LLMs that added Jev: across 80 games against a random player and Komodo Dragon it made zero illegal moves, winning 8 and drawing 22. | Benchmarks & evaluationMore guides & websites | |
| Minutes Jev voice evals — appuseminutes.app | — | Synthetic qualification script in the Minutes meeting-memory app that tests Jev on seven Choice decisions its voice path needs, such as attendee constraints, semantic recall, verified pastes, stale targets and prompt injection. | Benchmarks & evaluationMore guides & websites | |
| Models watching modelswww.southbridge.ai | — | Flags risky calls across 220,000 agent tool calls, shows that encoding tricks slip past it, and finds worded scales beat 1-to-100 ratings. | Benchmarks & evaluationMore guides & websites | |
| OmniJev/awesome-jev websiteomnijev.github.io | — | 🔥🔥 Papers, open reproductions and independent evaluations behind System One models and Jev. | Benchmarks & evaluationMore guides & websites | |
| One judge call or twelve dimension scores?agentjournal.dev | — | Measures a single direct Jev question per row against 12-14 Jev-scored dimensions with locally fitted weights on three classification tasks, using 34.1M input tokens for $1.43. | Benchmarks & evaluationMore guides & websites | |
| openjev-experimentsdemensdeum.com | — | Openjev experiments. | Benchmarks & evaluationMore guides & websites | |
| openjev-sglangekzhang--openjev-sglang-openjev.us-west.modal.direct | — | Jev-compatible API endpoint based on open models (prefill-only). | Benchmarks & evaluationMore guides & websites | |
| Qwen-2.5-1B-RLCD — demohuggingface.co | — | Parallel constrained decoding engine for Apple Silicon that answers multi-field decision schemas in one pass over stock Qwen2.5-1.5B, shipped as code only with no trained RLCD weights. | Benchmarks & evaluationMore guides & websites | |
| Rebuilding Nym's agent around Jevusenym.com | — | Replaces seven LLM reviewers and tool selection with Jev, falling back to a larger model when unsure, for a 4.1 to 5.7x speedup. | Benchmarks & evaluationMore guides & websites | |
| Rebuilding Nym's agent around Jev — appusenym.com | — | Replaces seven LLM reviewers and tool selection with Jev, falling back to a larger model when unsure, for a 4.1 to 5.7x speedup. | Benchmarks & evaluationMore guides & websites | |
| Security Incidents workflow evalevals.typesafe.ai | — | Official workflow eval for deciding whether a security alert is closed, escalated or contained; Jev scores 61.7% at $0.0001 and 0.3 s per case versus Opus 5's 66.2% at $0.0574 and 15.1 s. | Benchmarks & evaluationMore guides & websites | |
| SemIf (formerly OpenJev) — appopenjev.com | — | Independent Jev-like project that answers semantic if-questions by reading option logits from a frozen 4B open model on a home RTX 3090, with a browser demo and no text generation. | Benchmarks & evaluationMore guides & websites | |
| Six things I tried with Jevisaacflath.com | — | Six experiments swapping Gemini for Jev, from script fact-checking (24/24 correct, 0.41 s median) and news-feed ranking to PDF passage reranking (right passage first 7 times vs once) and agent-failure review. | Benchmarks & evaluationMore guides & websites | |
| System One models in high-throughput data pipelineswww.southbridge.ai | — | Entity resolution with Jev doing the bulk work and an LLM reviewing, at 226x lower cost; rewriting criteria as "what counts as sufficient evidence" fixed the dev set. | Benchmarks & evaluationMore guides & websites | |
| Typed Decisions benchmark — articlelatentnode.pages.dev | — | New 400-case benchmark for typed decisions where Jev scores 0.727, near the 0.735 teacher ceiling, against 0.646 for a frozen 149M ModernBERT encoder with small heads. | Benchmarks & evaluationMore guides & websites | |
| typesafe-vs-deepseek — apptypesafe-vs-deepseek.vercel.app | — | Side-by-side comparison of Jev and DeepSeek flash on speed, tokens, cost and accuracy across invoice extraction, email classification and reranking, plus fraud, guardrail and reconciliation pipelines. | Benchmarks & evaluationMore guides & websites | |
| WindTunnelwebmcp.com | — | WebMCP browser-agent benchmark of 49 tasks on 8 real sites across 21 configurations, where Jev + Mercury 2.5 tops the composite score, solving 49/49 tasks at $0.0011 median cost per task. | Benchmarks & evaluationMore guides & websites |
No matches this time
Try a broader search or reset your filters to explore the full collection.
Showing 917 of 917 resources