/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

TypeSafe AI debuts Jev, a model using “Reinforcement Learning for Calibrated Decisions” to return typed probabilistic decisions for use by software or AI models

‘Jev’ doesn't chat.  It produces typed probabilistic decisions  —  TypeSafe AI, a startup bestowed with $40 million in funding …

The Register Thomas Claburn

Context & Ripple Effects

TypeSafe AI’s $40M seed round was tied to a model designed to return numerical answers with probability estimates, positioning reliability measurement—not conversational generation—as its product focus. Jev makes that design concrete through typed, calibrated decisions that software and other AI models can consume.

The launch arrives as model makers also pursue reasoning capabilities, including Microsoft’s MAI-Thinking-1 reasoning model. It also speaks to the engineering concern raised by Andrew Barto and Richard Sutton that models should not be deployed without safeguards.

First-order effects

  • TypeSafe AI can sell Jev into workflows where an application needs a bounded decision and an accompanying confidence estimate rather than generated prose.
  • Developers integrating Jev can route typed outputs directly into software logic or downstream AI systems, making confidence thresholds an explicit part of the application design.

Second-order effects

  • General-purpose model providers face a sharper comparison on calibrated, machine-actionable outputs in decision workflows, rather than only on chat quality or broad reasoning performance.
  • Buyers building automated workflows gain a separate model category to evaluate for classification and routing tasks, increasing demand for testing whether reported probabilities match real-world reliability.

Third-order effects

  • If calibrated typed outputs prove dependable in production, AI application stacks may split between models that generate language and specialized models that make bounded decisions for software-controlled workflows.
  • The competitive frontier shifts toward operational assurance: training claims, evaluation methods, and interfaces that let customers set when an AI output is reliable enough to trigger action.

The trend: AI models are becoming more specialized for workflow execution, with calibrated machine-readable decisions emerging alongside conversational and reasoning models.

Discussion

  • @completeskeptic @completeskeptic on x
    We believe that the future is code + AI, so made workflow evals to reflect that Jev costs: $42 / BILLION input tokens ($0.042 / MTok) and output tokens are free (forever - they're too cheap to meter with our new architecture) Jev is named after Jevons paradox and off the intellig…
  • @completeskeptic Diogo Almeida on x
    After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I've spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x [vid…
  • @amasad Amjad Masad on x
    This is cool, but if your output domain is known in advance, why not just train a model to produce logprobs over enums?
  • @danshipper Dan Shipper on x
    we almost never test new foundation models but we've been testing this for ~a week @every and it's pretty wild. the kind of things that will be obviously indispensible in 6-12 months it doesn't produce words as output, it produces probabilities. so it can efficiently act as a jud…
  • @rohanpaul_ai Rohan Paul on x
    Another brilliant launch for developers: and its 20-200x faster than LLMs because it skips token-by-token generation entirely. TypeSafe AI just launched Jev, > 20-200x faster >40-400x cheaper (w/ output tokens free) > Frontier composable intelligence optimized for decisions So Je…
  • @k_grajeda Kevin Grajeda on x
    200× faster and 400× cheaper than llms this model is made for “decision-making
  • @completeskeptic Diogo Almeida on x
    We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI! ~10 calls/sec = ~$7/hour [video]
  • @completeskeptic Diogo Almeida on x
    Game: race from one Wikipedia page to another using only links Challenge: choosing between hundreds to thousands of links Shows not just intelligence-per-second, but also the compounding benefits of not hallucinating with high-cardinality choices [video]
  • @chrisgpt Chris on x
    I've genuinely never seen a model claim this low of a hallucination rate. This company was founded by Diogo Almeida, one of the researchers behind the instruction following work that led to ChatGPT, and instead of building an autoregressive chatbot that generates strings token by…
  • @noahpinion Noah Smith on x
    This is my friends' company. Pretty cool stuff.
  • @anatolikopadze Anatoli Kopadze on x
    i don't think people realize what just happened... 20-200x faster. 40-400x cheaper. output tokens free forever. the man who co-invented ChatGPT just dropped a new kind of AI called Jev and the twist is it can't write a word, it only makes decisions. he's calling it the shortest p…
  • @hammer_mt Mike Taylor on x
    A few weeks ago someone sent me a screenshot of the InstructGPT paper (that led to ChatGPT) with a name highlighted, and asked if I wanted to test a new language model Diogo was working on that doesn't output text... Couldn't resist, so here's my vibe check: https://every.to/...
  • @trikcode Wise on x
    Sir, he co-founded ChatGPT and now he's giving us intelligence for $0.042 per million input tokens with free output
  • @completeskeptic Diogo Almeida on x
    Extraordinary claims require extraordinary evidence so check out our release blog for more technical info: https://typesafe.ai/... Join our waitlist for early access: https://typesafe.ai/ Have technical chats and meme with us on Discord (rumors are good memers skip the line): htt…
  • @jxnlco Jason on x
    instructor v3
  • @completeskeptic Diogo Almeida on x
    The gains aren't free: Jev can't generate text Comparing Jev vs LLMs side-by-side makes the trade-off clear Fun fact: replacing sequential computation with parallel is the same way Transformers leapfrogged RNNs [video]
  • @timkellogg.me Mr. Tim on bluesky
    Jev: Fable-level model that doesn't charge for output tokens because they're too cheap to meter  —  it's not general though, it only makes decisions, doesn't generate text, but input tokens are measured by the billion ($42/btok)  —  typesafe.ai/blog/introdu...  [image]
  • @isolyth.dev Eris on bluesky
    This shit is fucking crazy - I have set Astra with the link on a journey to train our own.  If all goes well Erislab might shit out a vibed version of whatever the fuck they're doing, if they've make the mistakes of leaking enough bits for me to figure out what it is they are doi…
  • @scaling01 @scaling01 on x
    don't get one-shotted by this it's not a general language model and can't generate free form text it's probably a specialized diffusion model and it can only output a few different primitives and requires definitions of the output format
  • @willdepue Will Depue on x
    diogo is an immensely creative guy and is working on really different types of models with the principle of building composable, programmable AI systems from layers of small inferences. it's a weird and ambitious idea thats worth tinkering with
  • @benhylak Ben Hylak on x
    very impressed when i met diogo a year or so ago. i believe this is real.
  • @chaseleantj @chaseleantj on x
    This is a big deal. Right now, people use LLMs a lot as classifiers in production systems. But they're slow, expensive, and hallucinate. This model is supposed to be >100x faster, at $42/BILLION tokens, and generates not only type-safe structured output but also calibrated confid…
  • @hosseeb Haseeb Qureshi on x
    This is incredibly cool. Completely new form of AI models—output tokens are so cheap to meter, they're literally free. The deflation of intelligence continues. 👇
  • @davidsacks David Sacks on x
    Anthropic and OpenAI are already free to “pace the frontier” and should do so for business reasons, instead of first demanding a preferred regulatory framework
  • @stephenjudkins Stephen Judkins on bluesky
    This is extremely intriguing and might portend a future where LLMs perform many of the tasks they're currently good at vastly more efficiently and cheaply
  • @mergesort.me Joe Fabisevich on bluesky
    These claims are downright bonkers.  By *not* generating text, models become dramatically faster, cheaper, and more precise.  —  It's not an understatement to say that if this turns out to be true it would change so much about building with AI, and will lead to a dramatic spike i…
  • @timkellogg.me Mr. Tim on bluesky
    as far as i can tell, it's purely an RL feat  —  seems like they  —  1. take a fully pretrained LLM (decoder-only??)  —  2. slap a new 255-element classifier head on it  —  3. RL for Calibrated Decisions (RLDR)  —  i'm sure there's a strict format for the prompt, and during RL it…
  • r/singularity r on reddit
    TypeSafe AI releases AI model called Jev.  Rather than generating text, it makes decisions.  Its hallucination rate is far lower and its outputs are very cheap compared to traditional LLMs.
  • r/accelerate r on reddit
    “After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?  I've spent the last 2 years in stealth building …
  • r/singularity r on reddit
    New type of LLM released today “ it's optimized for structured outputs and can't hallucinate”