Skip to content

Repository files navigation

Awesome Jev Use Cases

English · 简体中文

Discover public projects using Jev for automation, model routing, search and business decisions. See what Jev judges, how software uses the answer, and what you can reuse. Use the Jev solution finder Skill to find relevant cases and plan your implementation.

Awesome Jev Use Cases — curated by SeeAPI

83 cases · Content updated: 2026-09-20

Use the Skill · Browse cases · Featured cases · Build with these cases · Suggest a case

Curated by SeeAPI, independently of TypeSafe. Cases link to original work; author results are not our measurements.

Find your solution with the Skill

Describe your task; let your agent find relevant Jev projects, compare their approaches and limits, and draft a workflow with source links.

With this repository available in your agent’s workspace, copy this request:

Read this repository’s SKILL.md. I want to route support tickets. Find 3 relevant cases, compare what Jev judges and what application code does, then propose a minimal workflow with a human-review path. Cite the original projects and distinguish documented behavior from tested results.

Get started / install · Read the Skill

Browse by use case

Category Cases What you can find
Content moderation & safety 8 Content screening and moderation
Automation & integrations 16 Browser, desktop, mobile and workflow tools
Model routing & code workflows 14 Model selection, code review and agent workflows
Semantic search & graph navigation 6 Search, reranking and graph navigation
Data classification & productivity 15 Tickets, spreadsheets and document analysis
Experiments & specialized applications 15 Games, simulation and specialized experiments
Benchmarks & behavior studies 9 Author evaluations and behavior studies

Featured cases

A selection across different tasks; placement is not a performance ranking.

Browser automation · Jev Ultrafast

A browser agent that turns visible controls into indexed candidates. Jev chooses an operation and target; a separate language model writes text when needed.

Reusable pattern: observed state → bounded action selection → execution → observation.

Jev primitives: Choice

Scope: the reported roughly 7.1-second flight search is a specific author-measured run, timed after the initial page observation. It is not a general browser-task speed guarantee.

Source · Full case

Flight-search result

Original author material: Flight-search result. Measurements shown are author-reported, not SeeAPI tests. MIT · Source · License and attribution

Model routing · Jev Codex Router

Classifies coding turns with Jev and applies a policy to select a model and reasoning depth, with logging and fallback behavior.

Reusable pattern: task classification → model selection → quality and cost evaluation.

Jev primitives: Choice

Scope: the reported roughly 60% savings comes from the author's replay of 237 real turns. It is not a SeeAPI measurement or a guaranteed saving.

Source · Full case · Try the recipe

Implementation walkthrough:Author's architecture and routing flow

Code review · Jev Review

Reviews diffs or codebases through staged judgments about risk, file profiles, evidence, mechanisms, severity, and conditional reviewer routing.

Reusable pattern: compose small judgments to focus deeper review on concrete regions.

Jev primitives: Choice, Noul, Score

Scope: an experiment that currently does not integrate compiler diagnostics or static analyzers. Findings are review leads, not proof of defects.

Source · Full case

Jev Review dashboard

Original material: Dev Agrawal · MIT · Unmodified · Source · License and attribution

Search reranking · Sift — search result reranking

A Chrome extension asks Jev about relevance, promotional content and depth, then reranks Google results in code.

Reusable pattern: Keep relevance judgments separate from ranking rules; users can restore the original order.

Jev primitives: Noul, Score

Scope: Judgments use result snippets rather than full pages; search intent affects filtering.

Source · Full case

Implementation walkthrough:Author's ranking and restore-order flow

Support ticket triage · JevTicketRouter — bilingual support triage

A .NET and React application classifies Persian or English support tickets. One request asks Choice questions for category and team, a Score for priority, and Noul questions for sensitive data and human review; local rules handle escalation and redaction.

Reusable pattern: Ticket → five typed judgments → local review and redaction rules.

Jev primitives: Choice, Score, Noul

Scope: The no-key demo can use deterministic mock answers. Low confidence forces human review rather than correcting category, team, or priority. Routing accuracy and redaction coverage were not tested.

Source · Full case

Author screenshot: security ticket held for human review in Mock mode

GhrezaKh74's application screenshot, explicitly in Mock mode with fictional tickets. The displayed decisions and probabilities are simulated, not Jev measurements or SeeAPI test results. Image loaded from the author's repository; its README declares MIT.

Original image · More author screenshots · License declaration

Community moderation · Jev Moderation Bot

A Discord moderation bot that evaluates messages and context for phishing, spam, and social engineering. Code applies escalating actions, and moderator corrections become safe precedents in later judgment context.

Reusable pattern: contextual message classification → moderation action → feedback.

Jev primitives: Noul

Scope: a text-message moderation project, not an image or video NSFW benchmark. False-positive performance has not been independently verified here.

Source · Full case

Implementation walkthrough:Author's moderation and feedback workflow

All cases

Purpose summaries and original sources for every project. Open a case for its implementation and limits.

Content moderation & safety

Safer with Jev — typesafe-on-neon — An HTTP gate that checks prompt injections, unsafe images, and unsafe replies before optionally forwarding a request. Its image pipeline first uses gemini-3-flash to describe the image, then asks Jev to judge risks in that description.
Reusable pattern: detection and interpretation → typed risk judgments → allow, review, or block.
Key limit: this is a vision-model-plus-Jev pipeline, not evidence of native image classification by Jev.
Source · Details & limits

Jev Moderation Bot — A Discord moderation bot that evaluates messages and context for phishing, spam, and social engineering. Code applies escalating actions, and moderator corrections become safe precedents in later judgment context.
Reusable pattern: contextual message classification → moderation action → feedback.
Key limit: a text-message moderation project, not an image or video NSFW benchmark.
Source · Details & limits

jev-spam-eval — Email classification experiments comparing natural-language Jev decision criteria with TF-IDF classifiers and combined scores, including tests across different mail sources and time periods.
Reusable pattern: written classification criteria → probability → threshold or ensemble.
Key limit: exploratory author-reported experiments.
Source · Details & limits

Jev MCP — jkudish — Exposes jev_verify, jev_screen, and jev_find to agents for evidence-based claim checking, input screening, and semantic candidate ranking.
Reusable pattern: package narrow judgments as reusable agent tools.
Key limit: verification depends on the supplied evidence; individual successful examples do not establish general accuracy.
Source · Details & limits

Capbroker — advisory screening around capability controls — A capability broker optionally uses Jev to flag suspicious MCP tool output and show risk advice at a human-approval prompt. Deterministic capability checks and the human approval decision remain separate from the model advice.
Reusable pattern: Capability enforcement → optional content warning or risk advice → human decision where required.
Key limit: The Jev layer is advisory and does not make the broker’s permission decision.
Source · Details & limits

Openroom — editable chat-moderation rules — The author describes a chat application using Jev, Convex, and Vercel to review messages before display. Natural-language room rules can be edited to trigger re-evaluation, with uncertain messages held for human review.
Reusable pattern: Message and room rules → moderation judgment → display or human-review queue.
Key limit: This entry is based on the author’s public description, not a verified backend implementation.
Source · Details & limits

JEVScan — Etherscan risk indicators — The author presents a Chrome extension that uses Jev to flag potentially malicious addresses and transactions on Etherscan. The collection records it as an author-demonstrated risk-indicator interface.
Reusable pattern: Blockchain-explorer context → risk judgment → on-page indicator for human inspection.
Key limit: The author’s post establishes the claimed integration, not its accuracy.
Source · Details & limits

BlueNoise — X reply noise filtering — A browser extension combines local keyword and account rules with an optional Jev noise assessment for replies that local rules do not match. Code uses the judgment to filter replies on X.
Reusable pattern: Local rules → remaining reply text and context → noise probability → display policy.
Key limit: The Jev option is experimental and disabled by default.
Source · Details & limits

Automation & integrations

Jev Ultrafast — A browser agent that turns visible controls into indexed candidates. Jev chooses an operation and target; a separate language model writes text when needed.
Reusable pattern: observed state → bounded action selection → execution → observation.
Key limit: the reported roughly 7.1-second flight search is a specific author-measured run, timed after the initial page observation.
Source · Details & limits

Jev desktop control in agent-desktop — A Jev integration reads an operating-system accessibility tree, selects an operation and target, and hands the result to a local desktop executor.
Reusable pattern: separate interface observation, model decisions, and execution.
Key limit: agent-desktop is a broader desktop tool with a specific Jev integration, not an exclusively Jev-based project.
Source · Details & limits

Typesafe MCP — itsmostafa — An MCP server connecting Claude Code, Claude Desktop, and Codex to Jev. Its evaluate tool accepts state and Noul, Choice, or Score questions for tasks such as ticket triage.
Reusable pattern: ask several independent typed questions about the same state.
Key limit: an integration tool; configurable examples should not all be counted as deployed customer use cases.
Source · Details & limits

SemDecide — Brings semantic predicates, routing, scoring, filtering, and guard decisions into Unix pipelines and CI through is, choose, score, filter, and guard commands.
Reusable pattern: typed judgments with explicit thresholds, uncertainty, and process exit codes.
Key limit: semantic judgments do not replace authorization or execution controls.
Source · Details & limits

typesafe-computer-use — A Mac automation loop converts screen information through OCR and deterministic processing, asks Jev to choose an action, and uses a writing model only when free text is needed.
Reusable pattern: Screen interpretation → bounded action selection → desktop execution.
Key limit: OCR and local processing provide perception; Jev does not directly inspect screenshots.
Source · Details & limits

Jev Browser — An agent skill and runtime that uses existing browser tools in a continuous observation, action, and verification loop. The planning agent supplies the goal and navigation guidance; Jev selects observed elements.
Reusable pattern: Plan once, then execute repeated bounded browser decisions.
Key limit: An unofficial integration requiring compatible browser tools; it is distinct from browser-use/jev-ultrafast and is not a browser service by itself.
Source · Details & limits

Mobile Jev — A mobile agent uses Jev to select actions on a real Android device through Mobilerun, with a studio, CLI, and execution traces.
Reusable pattern: Goal → mobile state → action selection → device execution.
Key limit: The documented Uber demo reaches payment selection, not a completed booking.
Source · Details & limits

zod-jev — semantic validation — Adds Jev semantic checks to Zod schemas, turning probabilities into validation issues alongside ordinary shape checks.
Key limit: Uncertain and unavailable judgments need explicit handling; a passing check does not establish factual correctness.
Source · Details & limits

HA-Jev — Home Assistant decisions — Turns home entity states into Jev probabilities, choices and scores exposed as sensors or automation responses.
Key limit: Device-state quality and automation policies determine whether a judgment is useful; integration was not run here.
Source · Details & limits

n8n TypeSafe community node — Exposes typed TypeSafe questions inside n8n so downstream workflow nodes can branch on answers.
Key limit: A community integration; installation eligibility and retry handling depend on the n8n environment and workflow.
Source · Details & limits

Unclutter — page clutter filtering — A browser extension classifies page elements with Jev and hides selected clutter using reusable rules.
Key limit: Uncertain elements should remain; hiding a consent dialog does not make a consent choice for the user.
Source · Details & limits

jev-mobile — Android Settings PoC — Uses semantic UI state, stability checks and bounded action choices for an Android observe–decide–act loop.
Key limit: The author exercised Android Settings; this is distinct from droidrun/mobile-jev and is not a general mobile agent.
Source · Details & limits

triage-guard — support, alert, and deployment judgments — A Python worked example shares a judgment engine across support tickets, operational alerts, and deployment risk. Batched Noul risk signals and a severity Score feed code-owned policy tables; the ticket flow adds department and urgency judgments.
Reusable pattern: Input → risk battery → policy thresholds → pass, review, block, or support route.
Key limit: The author labels it R&D.
Source · Details & limits

typesafe-jev-workflow — LangGraph email routing — An asynchronous LangGraph example sends email sender, subject, and body to a Choice question for invoice or general intent. Handlers set accounts_payable or general_inbox in graph state.
Reusable pattern: Mock email → Jev intent classification → graph branch and destination label.
Key limit: Handlers do not send email or make payments.
Source · Details & limits

Pi Jev Auto Mode — tool-call probability gate — A Pi extension combines deterministic command rules with Jev judgments before bash, write, and edit calls. Code compares condition probabilities with thresholds to produce allow, deny, or uncertain decisions.
Reusable pattern: Tool call → deterministic checks → semantic conditions → local execution gate.
Key limit: At review, README describes blocking uncertain results, while src/settings.ts sets uncertain to allow; src/jev/decide.ts treats an uncertain hazard-mode condition as satisfied.
Source · Details & limits

jev-skip — caption-based sponsor detection — A browser extension sends YouTube captions to Jev for sponsor-probability judgments on time segments. Local code displays a seek-bar heatmap and can skip selected segments without relying on a crowdsourced timestamp database.
Reusable pattern: Caption text → segment judgments → seek-bar overlay and optional skipping.
Key limit: No captions means no analysis; this is text classification, not audio or video understanding.
Source · Details & limits

Model routing & code workflows

Jev Codex Router — Classifies coding turns with Jev and applies a policy to select a model and reasoning depth, with logging and fallback behavior.
Reusable pattern: task classification → model selection → quality and cost evaluation.
Key limit: the reported roughly 60% savings comes from the author's replay of 237 real turns.
Source · Details & limits

Winnow — Judges blocks of long Claude Code tool outputs for task relevance. Confidently irrelevant blocks become summaries or stubs, while full text remains recoverable; uncertain blocks are retained.
Reusable pattern: reversible relevance filtering before context ingestion.
Key limit: judgment and summary generation are separate stages.
Source · Details & limits

Jev Review — Reviews diffs or codebases through staged judgments about risk, file profiles, evidence, mechanisms, severity, and conditional reviewer routing.
Reusable pattern: compose small judgments to focus deeper review on concrete regions.
Key limit: an experiment that currently does not integrate compiler diagnostics or static analyzers.
Source · Details & limits

jev-router — gargpratyush — Routes fresh user turns in Claude Code and Codex to fast or strong model tiers while launching the original CLIs.
Reusable pattern: Classify a turn and select a model without replacing the CLI.
Key limit: A separate project from Jev Codex Router by 0xNatoshi.
Source · Details & limits

eve — typed evaluation and model selection — The agent framework uses Jev by default for automatic model selection and typed evaluations; its documented tool-approval integration can escalate uncertain or failed reviews to a human.
Reusable pattern: Embed typed evaluation in model routing, tools, and approval decisions.
Key limit: Jev is the evaluator, not the sole model powering eve.
Source · Details & limits

DSPy typesafeify — hybrid inference — A proof-of-concept decorator routes boolean, enum and configured score fields to Jev while a generative model handles free text.
Key limit: The author's comparison uses three examples; it does not establish general cost or speed improvements.
Source · Details & limits

jevlogs — log triage — Scores diagnostic value and priority of OpenTelemetry logs before expensive analysis, while retaining an archive.
Key limit: Annotation alone does not skip downstream analysis; savings depend on forwarding mode and policy.
Source · Details & limits

SwarmRouter — agent assignment — Selects a specialist agent and judges ambiguity or the need for collaboration using typed questions.
Key limit: Routing recommendations do not demonstrate an executed multi-agent workflow; keyless demo responses are local heuristics.
Source · Details & limits

jev-axi — judgment CLI for agents — Provides typed commands for guard checks, build-log triage, diff review and bulk filtering, with reusable question recipes.
Key limit: The author's agent experiment reduced file reads without reducing cost; judgments do not replace source inspection or a complete safety boundary.
Source · Details & limits

Pi Warden — coding-agent rule feedback — A Pi extension checks edits against project rules and places feedback in the agent’s context. Other guards assess task drift, unsupported completion claims, and risky actions; deterministic patterns and model judgments feed code-owned responses.
Reusable pattern: Agent context and proposed changes → rule/risk checks → feedback or selected holds.
Key limit: Many findings steer or warn rather than block.
Source · Details & limits

commit-miner — commit classification and security-fix signals — A Rust CLI asks Noul questions about Git changes to identify bug-fix, security-fix, change-type, and CWE signals. Large inputs are reviewed in sections before selected evidence is used for a final judgment; local thresholds assign labels.
Reusable pattern: Commit diff → section judgments → selected evidence review → thresholded labels.
Key limit: Labels are model signals, not confirmed vulnerabilities.
Source · Details & limits

Foreman — semantic supervision of coding processes — An experimental runtime sends bounded task, worker-output, diff, and verification observations to nine Noul questions in one request. A deterministic policy uses those assessments to continue, start, stop, retry, verify, finish, or escalate managed work; assessments are printed for the CLI user.
Reusable pattern: Bounded worker observations → semantic assessment → local policy → process lifecycle action.
Key limit: At the linked commit, the worker interface exposes run and terminate, not a text-steering channel into a running agent.
Source · Details & limits

jev-belay — completion checks for Claude Code — A Claude Code Stop hook inspects the current turn’s transcript for file changes and verification evidence. When changes lack a subsequent passing check, it asks Jev four questions about the closing message; local thresholds and repetition limits determine whether to allow the stop or return feedback.
Reusable pattern: Local transcript evidence → conditional Jev judgment → allow stop or request follow-up.
Key limit: Errors fail open.
Source · Details & limits

jev-commit — commit-message and diff checks — A commit-msg hook, installable through the pre-commit framework, sends the staged diff and commit message to Jev for judgments about message quality, consistency, debug leftovers, unmentioned work, and credential-like content. Code applies thresholds and a separate credential check.
Reusable pattern: Staged diff and message → typed judgments and credential checks → local warning or blocking policy.
Key limit: By default, non-secret findings warn while likely credentials can block; strict mode also blocks other findings.
Source · Details & limits

Semantic search & graph navigation

Blink — Finds files from natural-language queries by having Jev score file and folder names, allocating walkers along likely paths.
Reusable pattern: narrow a search space through repeated semantic choices.
Key limit: result percentages represent the share of walkers reaching a file, not file correctness probabilities.
Source · Details & limits

neo4jev — Navigates a Neo4j graph by presenting outgoing relationships as Choice options, asking a Noul goal-completion question, and exploring candidate paths with beam search.
Reusable pattern: model-guided edge selection inside a deterministic search algorithm.
Key limit: a demo with explicitly labeled stand-in answers when real TypeSafe calls fail.
Source · Details & limits

Sift — search result reranking — A Chrome extension asks Jev about relevance, promotional content and depth, then reranks Google results in code.
Key limit: Judgments use result snippets rather than full pages; search intent affects filtering.
Source · Details & limits

Every — function-level semantic search — Parses source into functions and asks Jev a yes/no question for each, returning ranked matches with cached scores.
Key limit: Function-local judgments do not establish whole-program dataflow; scanned source is sent to TypeSafe.
Source · Details & limits

Jev Search — intent selection and result reranking — A TypeScript application asks Jev to select search sources, time ranges, and query candidates, fetches results through Search1API, then judges title/snippet relevance in batches. Code merges URLs and ranks results by relevance, engine agreement, and original position.
Reusable pattern: Search intent → external retrieval → per-result judgments → merged, streamed rankings.
Key limit: Relevance scores do not verify page facts; snippets can be incomplete or stale.
Source · Details & limits

jev.nvim — semantic function search in Neovim — A Neovim plugin uses Treesitter to split buffer code into functions and asks Jev whether each function matches a natural-language question. Results appear as probabilities in virtual text and a ranked quickfix list, integrating semantic search into the editor.
Reusable pattern: Buffer or selected files → function extraction → per-function judgments → ranked editor results.
Key limit: Functions are judged separately, without cross-function context; matches are search leads rather than confirmed defects.
Source · Details & limits

Data classification & productivity

Judge Sheets — predictive spreadsheets — Typing a column header such as Urgency lets Jev infer a prediction schema; confirming it fills rows through JUDGE, PICK, and RATE functions, with grouped requests and streamed updates.
Reusable pattern: Header intent → typed schema → row judgments → spreadsheet recalculation.
Key limit: A standalone spreadsheet demo, not a Google Sheets integration.
Source · Details & limits

Notra — typed evaluation in analytics — The codebase includes a Jev evaluation client through Vercel AI Gateway, a NOTRA_JEV_CLASSIFIERS flag, and optional typed evaluation alongside brand-mention analysis.
Reusable pattern: Introduce typed judgments into an existing analytics workflow with an LLM fallback.
Key limit: Source inspection establishes an integration path, not independently verified production deployment or latency.
Source · Details & limits

human-compiler — writing diagnostics — Combines local text analysis with Jev questions about clarity, intent and tone; deterministic rules render diagnostics.
Key limit: Diagnostics reflect chosen rubrics and thresholds, not objective writing quality or generated explanations.
Source · Details & limits

Kill My Idea — idea scoring — Jev scores a product idea against several criteria; local weighting maps the results to a product verdict.
Key limit: Heuristic feedback, not validated prediction of business success; the project also supports mock data.
Source · Details & limits

Jev CV Screening — Stores typed CV judgments separately from local scoring rules, allowing supported policy changes to reuse existing answers.
Key limit: New questions require new judgments.
Source · Details & limits

Jevibe Check — social tone labels — Labels Bluesky posts and drafts using Jev choices, with custom classifiers and filtering controls.
Key limit: Text-only analysis excludes images, videos and wider conversation context; sarcasm can be misclassified.
Source · Details & limits

JEV Resume Analyzer — Reviews extracted CV text against explicit rubrics and optional job requirements, showing questions and answer distributions.
Key limit: Missing, inapplicable and unassessable evidence remain distinct; it does not provide a validated hiring prediction or ATS score.
Source · Details & limits

LaneBreak — support ticket routing — Uses Choice for team assignment, Score for priority, and Noul for refund intent, churn signals and human escalation.
Key limit: Without an API key the implementation uses local heuristic demo responses; those are not Jev results.
Source · Details & limits

Jev Column Race — review annotation — Batches sentiment, topic, bug and churn judgments over app reviews, then allows local reranking; includes a comparison with Gemini.
Key limit: The published timing is from a recorded run pair.
Source · Details & limits

JevTicketRouter — bilingual support triage — A .NET and React application classifies Persian or English support tickets. One request asks Choice questions for category and team, a Score for priority, and Noul questions for sensitive data and human review; local rules handle escalation and redaction.
Reusable pattern: Ticket → five typed judgments → local review and redaction rules.
Key limit: The no-key demo can use deterministic mock answers.
Source · Details & limits

Transcript Scorecard — incremental call evaluation — A proof of concept replays fictional support-call transcripts incrementally. Each enabled criterion contributes a Score and a Choice selecting an evidence sentence; code normalizes and weights the results, then stores the final evaluation in SQLite.
Reusable pattern: Growing transcript → criterion scores and evidence selection → weighted score history.
Key limit: This is transcript replay, not verified live audio recognition.
Source · Details & limits

Paper Trellis Citation Verifier — citation support review — A manuscript-review tool pairs citing sentences with source passages. Claude can locate quotations, code checks quotation presence, and Jev chooses supports, contradicts, or says_nothing for the sentence and selected passage; the reviewer retains the final decision.
Reusable pattern: Citation matching → passage selection → three-way support judgment → human review.
Key limit: Jev reads a bounded passage around a quotation, or the source opening, rather than the entire paper.
Source · Details & limits

Research Desk — staged news and company judgments — A demonstration uses company profiles and headlines from yfinance in a staged Jev pipeline for relevance filtering, ranking, and mechanism matching. A request view exposes the state and typed questions behind the displayed judgments.
Reusable pattern: Company/news inputs → staged judgments → code-based filtering and traceable results.
Key limit: The author describes thresholds as initial guesses rather than values fitted to outcomes.
Source · Details & limits

JevFilterForX — timeline value scoring — An X extension asks Jev to score signal, actionability, and originality. Local weighting produces a 0–100 value score, while topic and noise labels support filtering and focus modes.
Reusable pattern: Post text → rubric scores and labels → weighted score → timeline filtering.
Key limit: This is timeline scoring, distinct from BlueNoise’s reply filtering and Jevibe Check’s Bluesky labels.
Source · Details & limits

JevScout — career-page navigation and job matching — A coding-agent skill starts from a company website, uses Chrome through CDP to observe and navigate pages, and asks Jev to judge links and rank jobs against a job-seeker profile.
Reusable pattern: Company website → career-link judgments → job discovery → profile-based ranking.
Key limit: A demo MVP with mock and fixture-based workflows.
Source · Details & limits

Experiments & specialized applications

TypeSafe AI Playground — A Rust CLI exploring tasks such as protected health information detection and code-comment review through typed questions and scores.
Reusable pattern: reuse decision primitives across clearly defined application criteria.
Key limit: experimental tooling, not a privacy-compliance certification.
Source · Details & limits

Prism's Jev judgment service — Maps liquidity-strategy questions about distribution choice, toxic flow, recovery holding, and market stress to Choice and Noul judgments alongside existing heuristics.
Reusable pattern: compare model advice with existing rules in shadow or advisory mode.
Key limit: the inspected Jev module explicitly states shadow/advisory use.
Source · Details & limits

1v1 Jev — Quickscope Arena — A browser FPS opponent controlled through Choice/Noul questions about movement, aim, ADS, firing, and jumping. The server supplies structured game state and includes a heuristic fallback.
Reusable pattern: repeated bounded decisions drive a real-time interactive agent.
Key limit: the README's approximately 9 Hz loop describes this project, not a universal Jev performance figure.
Source · Details & limits

jev-trader — A trading experiment asks Jev for buy/sell judgments from the Kuru MON-USDC order book on Monad, with code handling quotes, limits, and execution.
Reusable pattern: Order-book state → directional judgment → program-controlled order handling.
Key limit: The default model is a mock momentum heuristic; Jev requires explicit configuration.
Source · Details & limits

TypeSafe Mario — An emulator harness converts telemetry and RAM into structured state; Jev selects NES controller actions and provides jump and danger judgments.
Reusable pattern: Structured game state → Choice/Noul/Score → controller input.
Key limit: The model does not receive screenshots.
Source · Details & limits

jev-drone — A MuJoCo quadrotor simulation converts camera depth and segmentation into symbolic scene data; Jev advises maneuvers and risk while conventional code handles flight control and safety.
Reusable pattern: Perception in code → tactical judgment → guarded control.
Key limit: Simulation rather than real-world flight; Jev receives JSON rather than images and is advisory.
Source · Details & limits

tsai-sc — StarCraft Strongarm — A harness reads structured game state, asks Jev to choose commands, and executes mouse and keyboard actions in the original StarCraft shareware mission Strongarm.
Reusable pattern: Structured strategy-game state → command selection → input execution.
Key limit: The game pauses during state reads and inference.
Source · Details & limits

HEIST ONE — stealth-game guards — Jev judges threats, suspicion and intentions for guards; server code controls legal actions, physics and fallbacks.
Key limit: Scripted mode is available; a recorded live run does not establish repeated success.
Source · Details & limits

TypeSafe Minecraft — structured action control — Jev selects Minecraft actions from structured observations; Mineflayer executes them with code-supplied candidates and checks.
Key limit: The earlier video used high-level control; the newer direct-action controller is a separate experiment, not screenshot-based vision.
Source · Details & limits

Jev for Engineers — Eight Python examples apply typed judgments to engineering workflows such as CAD, BOMs and change control, with decisions made in code.
Key limit: Small synthetic examples and uncalibrated thresholds do not establish suitability for real engineering decisions.
Source · Details & limits

Jev literature screening — Combines inclusion choices and atomic eligibility judgments for title-and-abstract screening against an author's documented review protocol.
Key limit: The current README evaluates Cohen ADHD abstract triage; limited abstracts and filtering rules can miss eligible papers.
Source · Details & limits

Jev JFK Simulation — voice-driven airport demo — An author demonstration combines a simulated JFK airport with real-time voice models for radio interaction and Jev for operational judgments. It illustrates separating voice interaction from a bounded decision loop.
Reusable pattern: Simulated airport state and radio interaction → Jev judgment → simulated response.
Key limit: Evidence is the author’s public post and demonstration; no public implementation or complete request trace was verified.
Source · Details & limits

Jev Canvas — voice and gesture canvas demo — Jack Cheng’s author demo combines voice, pointing, and canvas state to present Jev action and target judgments for manipulating shapes. It is a creative-tool interaction example rather than a general image-generation model.
Reusable pattern: Voice/pointing inputs and canvas objects → action and target judgment → canvas operation.
Key limit: Evidence is the author demo recorded in the collection; the original post was re-opened, but this pass did not independently replay the full video.
Source · Details & limits

Jev Gomoku — candidate-move selection — A nine-board Gomoku experiment supplies textual board state and code-generated candidate moves to Jev. Each move uses a Choice question; five input configurations explore the effect of tactical facts, coordinates, directional lines, and short lookahead.
Reusable pattern: Board state → deterministic candidate generation → Choice → legal move execution.
Key limit: Jev selects from heuristic candidates rather than all empty squares.
Source · Details & limits

jev-plays-pokemon-red — bounded game decisions on PyBoy — A Pokémon Red experiment reads emulator RAM into structured state. Deterministic Python handles routes, battle arithmetic, and legal actions; Jev selects among candidates at branch points. The harness records turn-level faint predictions and outcomes for later Brier-score evaluation.
Reusable pattern: RAM-derived state and legal candidates → branch-point judgment → emulator action and outcome recording.
Key limit: This is a code-guided experiment, not autonomous long-horizon planning or screenshot-based play.
Source · Details & limits

Benchmarks & behavior studies

jev-sec-bench — security judgments — Evaluates prompt-injection detection and vulnerable-code judgments with published datasets and per-sample results.
Key limit: Reported results depend on context and thresholds; this is not an NSFW benchmark or a complete security boundary.
Source · Details & limits

Jev Behavior Study — Studies question framing and failure modes through text tasks, Snake and a 3D city, with reports and recorded traces.
Key limit: Synthetic task-specific observations; repeated calls are not independent problems, and assisted control differs from direct control.
Source · Details & limits

jev-rerank-bench — retrieval evaluation — Compares Jev relevance rubrics with other rerankers on shared BM25 candidates and publishes saved responses and scoring code.
Key limit: The headline averages do not establish a winner; weighting datasets versus queries changes the comparison.
Source · Details & limits

jev-phishing-bench — email signals — Compares direct phishing judgments with atomic Jev signals combined by a local classifier.
Key limit: Synthetic emails contain potential shortcuts; direct and held-out experiments have different test sets and must not be conflated.
Source · Details & limits

jev-headline-bench — headline selection — Asks Jev to choose between historical Upworthy headlines and compares choices with recorded click outcomes.
Key limit: Historical within-article pairs do not replace a randomized A/B test on a new site's audience.
Source · Details & limits

Jev judicial-text annotation — Compares typed annotation of 12 variables in 120 Portuguese judicial documents with generative-model structured outputs.
Key limit: The reference process includes model-generated labels and adjudication; reported accuracy is not based entirely on human gold labels.
Source · Details & limits

LLM Chess Jev Player — constrained chess evaluation — An adapter adds Jev to an existing chess evaluation framework. For each move, application code provides the FEN position, side to move, and legal UCI candidates; a Choice answer is converted into a make_move action.
Reusable pattern: Board state and legal moves → Choice → move execution and game records.
Key limit: Legal candidates are supplied by code, so protocol success does not establish chess strength.
Source · Details & limits

Every Judgment Lab — writing and knowledge-work checks — Mike Taylor’s experiment suite breaks writing review, context retrieval, and business triage into bounded judgments. Its writing experiment evaluates 37 documents against 21 criteria; the collection counts the suite once rather than treating its 11 experiments as separate projects.
Reusable pattern: Documents or task state → parallel rubric judgments → flags for further review.
Key limit: The reported 777 judgments in under 0.7 seconds are author measurements, not our benchmark.
Source · Details & limits

Jev Maze Lookahead — a negative planning experiment — A maze experiment compares parallel future-step questions with single-step decisions and explicit adjacent-tile hints. The project separates a deterministic BFS mock from real-API runs.
Reusable pattern: Maze state → proposed moves → environment checks → recorded successes and failures.
Key limit: The author reports zero solved mazes in the quick multi-step setting; with adjacency hints and one next-move question, 6 of 10 small 5×5 mazes were solved.
Source · Details & limits

Build with these cases

What is Jev?

Jev is TypeSafe AI’s System One model for typed judgments. Choice selects candidates, Noul evaluates a yes/no proposition, and Score rates ordered criteria. Application code decides what to do with those answers. Model background.

Model origin & access options

Start with the official documentation or the provider and SDK guide. This collection does not assume provider interfaces are interchangeable.

Evidence & scope

Read evidence definitions and the full review scope. This release adds offline recipe checks, not independent model benchmark results.

Sources & contributions

Suggest or correct a case · Source history · Maintenance guide · Changes.

License

Original documentation: CC BY 4.0. Code: MIT. Third-party materials retain their own rights and attribution; see scope.

TypeSafe official direct access · Third-party platform access · Agent tools & MCP integrations · How access resources are selected · Three typical judgments · Python example: ticket classification

About

Explore real-world use cases and projects built with TypeSafe AI's Jev: content moderation, AI agents, model routing, and semantic search. Curated by SeeAPI.

Topics

Resources

Contributing

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages