The Plug Was Wrong: Why TypeSafe Built Jev for Code, Not Chat

Latent Space — Diogo Almeida, TypeSafe co-founder and CEO, with swyx — “Why We Made Jev,” cFx9Z3ZXca0


Diogo Almeida’s question is not whether models are smart. They are smart enough to threaten millennium-prize mathematics and still fail at rote work that nobody enjoys doing. The engine, he says, is supercharged. The plugs are wrong. Economically valuable software cannot consume a chatbot, a refusal-prone coworker, or a reasoning model that was optimized to look good on verifiable benchmarks. It needs intelligence that code can call the way it calls a database: calibrated, typed, cheap, and willing to answer.

That is the product he is selling. TypeSafe launched Jev in the days before Almeida sat down with swyx, and Almeida is not a disinterested critic of the stack he left. He is the technical CEO of the company that trained a new class of model for that plug. The argument still has to be taken on its own terms, because it is more specific than “AI should be in software.” It is a claim about the task post-training optimized for, and about why instruction-following and RLVR both pointed the field at the wrong customer.

A ragged corpse, briefly in sync with reality

Emotionally, he said, he had never been worse: a ragged corpse fighting fires. Mentally, for one week, the carnival house of mirrors that is the AI field felt aligned with reality. Developers got it. An AI-based economic revolution was “back on the table.” He had skipped VIP investor conversations to prioritize a Discord town hall. The Discord, swyx noted, was now 100,000 people. Almeida would rather talk to that room than to “really important people” he would not name. In his ideal calendar, community is the default, not an afterthought he squeezes in while walking to a studio.

He is hiring marketers because they do not have one. People called the launch marketing genius; he called it being goofy and irreverent, offboarding a waitlist as hard as they could. Waitlist signups, he said, do not matter for a developer platform. A large share of them are not developers. They arrive looking for ChatGPT 2.0, try a few queries, and bounce. One power user’s for-loop, running in the background as a dependency, swamps every human on earth writing a couple of queries. The milestone he would claim is not signups. It is tokens per day that keep churning at night — machines calling, not people clicking around. He said they had surpassed a trillion tokens a day. That figure is his, from a week-old launch, relayed while he admitted he is not on top of the dashboards.

The launch video had something like 36 million views, rounding toward 38 while they talked. He would rather wear the parody swag — “your favorite neo-lab’s favorite neo-lab,” “neo-lab with product” — than be filed as a neo-lab. What he wants is a reliable developer platform. Platform uptime, he claimed, had more nines than Anthropic during the most unprecedented launch they had ever had. That, too, is founder-reported.

System One, because the consumer is code

Jev is TypeSafe’s first large programmable model. The most accurate name they have for the class is System One models: machine-native, meant for code to consume. They are not attached to the branding. They are attached to the customer. Pretrained LLMs autocomplete the internet. RLHF models reply to text. RLVR sits in a gray area next to RLHF. Jev is supposed to emit things programs can take as inputs. Hence TypeSafe.

They do not call them “decision models,” even though the current primitives are decisions, because Almeida says System One goes beyond that and they have “stuff in the tank.” The launch was not supposed to be the big one. swyx suggested he should have labeled it a low-key research preview. Almeida said it kind of was.

The name is Jevons: intelligence per dollar. He likes arguing about the relative importance of reliability, cost, calibration, and speed. Jev, as a series, is supposed to sit on the frontier of intelligence per dollar. Intelligence per second is a different metric — real-time and user-facing budgets between 100 milliseconds and a second, where halving latency can mean twice the sequential intelligence — and he does not think that will be Jev’s niche. Cerebras, Etched, 100,000× speedups: interesting, not the brand. Internally he hunts people down if a smarter model leaves the pre-frontier of intelligence per dollar. Holding intelligence constant while getting faster and cheaper is, he agreed, the load-bearing claim, and the hard one.

Why long strings look doomed, and why they aren’t

Yann LeCun’s infamous slide — language models are doomed because error probability compounds with sequence length — is, Almeida said, mathematically obvious and empirically wrong. That is his favorite teaching example of a disconnect. A calibrated, mode-covering distribution is not overly punished for outliers. You expect to be out of distribution some of the time. Pre-GAN image models made blurry pictures because they covered the distribution. GANs mode-drop: they abandon minority classes and emit the common ones. RLHF does a version of that to language. To generate long strings without visible errors, the model becomes extremely conservative. Errors are easy to see. Subtle, confident-looking mistakes are not. Calibration, in that regime, is “total poison” in the probability distribution of strings. Overloading string models as decision engines is a bad time.

That is the downside of RLHF almost nobody noticed in the launch materials: mode dropping, which he treats as the same phenomenon as mode collapse. Chat-tuned models collapse toward what you want to hear, or what is most likely, rather than their internal confidence. Latent Space had already covered calibration with Hugging Face; Almeida treats that as nearby evidence. Almeida wants a blog post. For now the spicy take is that LeCun is among the most accurate public critics, and still wrong about this particular pie chart.

JEPA, world models, joint-embedding prediction: cool early research. He would not say whether it is practical. Scaling laws tell you how much better you get per unit in, usually sublinear gains for exponential resources, which is a bad investment unless those gains are extremely valuable. He would rather talk about his bitterest lesson.

The bitterest lesson is the task

Sutton, roughly: algorithms beat compute. Almeida’s version is harsher. Data matters more than compute. The right task — a north star — is the hardest and most important thing. In LLM land this has happened “twice so far, maybe 2.2 times.” RLHF shifted the task to instruction following; nobody realized that was possible. RLVR was a tiny edit to the direction. TypeSafe’s RLCD is a new task: programs in the loop.

He has not published a paper on RLCD. swyx asked him to prove it is not jargon. Almeida’s move was to undefine RLHF first. The original Christiano-era work — he recalled a robot backflip, PPO, hard-to-specify outputs — is one meaning. Learning to summarize, PPO on language models, InstructGPT’s authors, is another. What he means by RLHF is not the algorithm. DPO and its descendants still “do RLHF.” The thing that matters is the north star: instruction following. RLCD is the analogous move for programmable AI. Remove the human from the loop of the thing you are tuning, because the customer of the intelligence is going to be a program.

They think of themselves as a data lab, not a model lab. Model capabilities means data. Data is what gets nines. They are hiring “infinite” data people. Synthetic data is only the surface; the shape of the task changes the shape of the data. RLVR’s data looks like environments. RLHF’s data is human feedback. Theirs is its own kind. They do not want to train on user data even if they could get the terms, because real-world queries follow a power law, overfit, and fracture the model toward the present. They are aiming at sci-fi infrastructure years down the stack — he analogized pretrained LLMs to UDP and their models to TCP — and present-day logs would overfit the present. The data people study a “cognitive core” he claims is less jagged than anyone else’s, find the jaggedness, and address it surgically in the general case. He onboarded them with a talk he assumed would run longer than a two-hour interview.

They could have released Jev a year and a half earlier “if we wanted to be dumb.” Last year’s fundraise hurt because they would not benchmax; nobody believed them. They stood on it anyway. Public benchmarks, he said, are antithetical to selling a je ne sais quoi of intelligence. They are extremely gameable. Every lab used to have a team collecting MMLU-shaped data. Private proxy benchmarks he is “medium” on. What he wants, long-run, is vibes and trust until a buyer puts the model in a real workflow and measures that workflow. Internal evals exist. He rules “not bullshitting ourselves about how smart the model is” with what coworkers might call an iron fist. The preview-period terms that looked like they banned benchmarking are, he said, coming out; they are not trying to stop people. They are trying not to let a public leaderboard become the loss function.

A refusal is a type error

Discord keeps asking why TypeSafe is “opposed to safety alignment” and does not refuse. He is not opposed to safety as a principle. He thinks safety alignment is generally misaligned with users, and that refusal is a type error. In a chat product, “I’m sorry I can’t read DNA.py” is annoying, and Stockholm syndrome makes you work around it. Buried in a dependency, stochastic refusal is insanity. Someone else’s software receives a weird message and the whole program breaks. That design, he said, comes from people who do not understand software and are obsessed with the horseless carriage of an “AI coworker” instead of the full power of AI.

Capability alignment is doing what the user wants — good for software engineers, more nines, eventually as unthinking as a database query. Safety alignment is the opposite of instruction following: following someone else’s instructions, OpenAI’s or Anthropic’s. Fine for ChatGPT if they do not want NSFW roleplay because parents are in the user base. Nuts in an API. He would prefer their models not be used to kill people. He will put a thumb on the scale socially. He will not do it at the technological layer, because every extra overfit fractures intelligence. Databases do not get to decide whether the CIA may query them. When companies Slack to ask if they can deploy, his answer is: you are a developer, it is none of my business, and if the task is properly decomposed TypeSafe should not even be able to see the downstream use. Bias comes out of the technological layer “as long as I’m in charge.”

swyx’s pushback was war. Almeida’s reply was that a general-purpose technology is the wrong place to encode that preference. The pacing-the-frontier conversation among researchers — slow down, the public is not ready; every frontier lab co-signing a pause document after the OpenAI “blip” — he reads as a sleight of hand. It assumes everyone must do more RLVR, more “do anything in the middle,” because that is how you make models powerful on hard problems. Sandboxing, he said, was an obvious problem they could have solved; they chose not to, because the more you let the model do anything in that middle stretch, the more powerful it becomes outside it. RLVR, he said, is not actually about verifiable rewards; that was failing before the reasoning revolution. Instruction following was the bastard child of early post-training; pre-training teams did not want to wait on human evals; codegen was burning resources on unit-test RL that needed reasoning to work. Zero RLVR is the optimal amount for TypeSafe’s shape. The only people at fault in the closed-minded version of this debate, he said, are the researchers. The public reasonably assumes OpenAI and Anthropic are doing the best they can. His goal is not to convert labs. It is to spark software engineers to automate what they have always wanted automated.

Reliability is robustness, not a seed

Big reasoning models are smart and still not reliable enough to replace an intern on work that is already economically worth automating, because they were optimized for something else. He wants to automate easy work before hard work. The dream is flow state: you need a non-trivial branch, you write a TypeSafe System One query, it branches accurately, you do not even try the model to know it will work. He does not know if “sufficiently reliable” exists. He knows it is a long slog.

Reliability, for him, is a catch-all for whatever stops automation: type safety, jaggedness, trust in outputs. Determinism — same inputs, same outputs — is slightly interesting for unit tests and the wrong north star. What he thinks people actually need is robustness: similar inputs, similar outputs. Their test is to drop UUIDs, “nonses,” into a prompt and demand similar answers because the question is semantically the same. That is where people get burned letting AI make decisions. Determinism can be offered; it costs intelligence per dollar. They are GPU-constrained and trying to start a California gold rush of weird experiments, so anything that burns more GPUs for the same intelligence is in tension with getting the model into as many hands as possible. swyx predicted peer pressure will force seeds anyway, as it did at OpenAI. Almeida is stubborn. This week he said his chief of staff, Kay, was the most powerful person in tech.

He will not silently change a deployed model. That is insane for an API, whatever first-party chat products do. They will launch new models faster than people are used to, and they are not promising long-term support, because they think there is a lot of improvement left. They might temporarily LTS Jev 1.13.0 if too many people depend on it; the alternative is fracturing the fleet into a hundred versions. Research is cooking on a “really sick” LTS story. It is not current models. They will get smarter. Between nearby versions, he claimed, deltas are often smaller than calling a string model twice; the big jumps are jagged-to-wow. Porting LTS models to other silicon: no comment. Quantizing launch quality away to free GPUs: swyx warned buyers would fear it. Almeida’s public line is they will not change deployed models, and they will not fly blind — they plot internal evals — but they will keep doing “absolutely disgusting things” on the Pareto frontier of intelligence per dollar. People told him not to call the stack a Frankenstein’s monster. He thinks Frankenstein’s monster was the innocent one. He has not read the book. Sleep is the priority.

Noul, choice, score — and stop stuffing the system message

Three primitives, all deliberately not existing types. Noul is bool-ish but continuous, named from Bernoulli, after internal debates about “peool,” “pool,” and a “pool party” nobody allowed. A bool would confuse people. Score is not an int or float; mapping it through Instructor or Pydantic as one will get you cooked. It is closest to LLM-as-judge. Choice is closest to a function call, which he called an extremely disgusting thing, and maps to a switch on an enum. Nouls map to if-statements. Scores map to sorting or thresholding. More types will come, still mapped to programming primitives. Internally almost everyone hated the name Jev; they have all apologized except one person who wanted it called Meow.

State, instructions, and criteria can all be structured JSON. Programs insert them. You do not template a system message. System messages are “disgusting global variables.” Passing numbers as strings is what you do when a human is printing. Inside the machine you want nested semantic structure. Every extra nested level is harder to reason about; they are cooking on that. Nobody had asked him for this kind of API guidance in months, he said — not since he was onboarding DevRel. His strongest usage advice is not a monetization trick, he insisted: ask lots of questions. Decompose into the smallest semantic unit. Make each decision evaluable. “Don’t read this subdirectory” or “don’t pass API keys to DeepSeek” should be programmatically close to guaranteed, not hoped-for inside a giant prompt that another giant model then grades. You will never have guarantees from ML. Breaking the task down is how you measure it, how you add a failing case as another question, how you pick a threshold from real examples, how the bug stays fixed instead of rotting out of a context window. “ML without the ML.”

swyx had already benchmarked the old way: one fat system prompt versus a hundred small calls. The decomposed version was slower, more expensive, and worse — which is why people stopped doing it. Almeida did not deny the inconvenience. He said it produces something you cannot rely on, and it would break his heart if their models could not run in the background. On refusals specifically: do not ask “should I refuse here.” Ask many independent questions about the situations that warrant refusal. If it fails because you did not specify a case, that is software engineering. Add the question, add the threshold, keep the test.

Calibration is not perfect. He did not claim it was. Fine-tuning is not offered. He can imagine it, fears it as a footgun — general models’ extra tasks can help edge cases; narrowing can hurt — and noted that OpenAI, Claude, and Gemini shipped fine-tuning and took it back, leaving it to open source. Cascades of model sizes, dynamic routing, even automatic fine-tuning on a Pareto frontier: sci-fi cooking, not a promise. Different sizes of Jev: absolutely, because he cannot know how much intelligence a given call needs. Demand is unlimited. People are already telling them not to ship more because it is good enough. He thinks that is lame. Culture is what you do when the market does not reward it. He hopes they become as boring as Visa. He also hopes they do not stop at the opening salvo.

An inverse SaaS-pocalypse, if the nines arrive

How can 2019 SaaS still look like 2019 SaaS in 2026, except for a chat box on the side that cannot be trusted with decisions the company has stakes in? That, to him, is nuts given the economic incentive. He wants AI to stop being the foreground character and disappear into software. The bet is an inverse SaaS-pocalypse: existing software companies, who already know what is worth automating, get supercharged. He has a line about TFP growth of 3% in five years. swyx had never seen a lab care about TFP. Almeida tied it to the old OpenAI charter — majority of economically valuable work — and to the contradiction he will not let go: models that do millennium-prize math and approximately zero of the world’s economically valuable work. All models are roughly tied at zero; maybe they have started, probably not 1% yet. When it happens it will show up in economic statistics. It will not, he claimed, cause mass unemployment. It will cause shifts, and the world will be better. That last cluster is aspiration, not evidence.

System One versus System Two is empirical, like scaling laws, like why robotics still does not work despite the spend. Pretrained condensations of intelligence are fundamentally System One thinkers. RLVR did incredible, fragile, fractal System Two work; he is in awe; he does not think it produces AI doom, and 0% would be miscalibrated. ChatGPT-era models were “general but bad at GSM8K.” RLVR models are jagged in the opposite way. North stars, as he tabulated them: RLHF is please humans; RLVR is optimized benchmarks, because a benchmark is programmatically verifiable by definition; RLCD is make it reliable for programmatic use. swyx threw Jev at tasks on day one: single-hop state of the art, never use anything else; multi-hop degrades monotonically with hops. Almeida’s reply is that they unearthed what the condensed cores already contained. System One is the description of what works, not a theology. He is not promising “no reasoning, ever.” He is promising a machine-native ROI north star. Some less slow, less fragile forms of reasoning are “on the cards.” Vision is on the cards. Latent/continuous reasoning inside the model is a terminology fight with swyx he did not want to turn into a product promise.

Pre-launch, more than half the people who played with it did not get it. Non-technical teammates were terrified they were selling a vitamin. Almost no revenue. Technical people saw computational properties “off the charts.” The people who did get it asked about procurement. So they targeted developers, bet on FOMO, and watched companies offer GPUs because rate limits, not signups, were the constraint. He wants to be held to launching things that are better for developers than enterprises. He dyed his hair to go talk to them during the company’s most important week because it felt dirty not to.

Use cases they mapped from first principles before release: dark data — piles of corpus too expensive to throw LLMs at, a data scientist’s dream, plus coding agents as the volume businesses; real-time loops, especially e-commerce and assistants, where CEOs know what 10 milliseconds are worth; “verify everything,” observability over other LLM calls, cheap parallel questions on one paid-for state, IDs on every message; “smart software,” including a programming-language-as-Jev project he found cooler than TypeSafe’s own demos. Computer use and voice control of a desktop arrived out of left field. He is anti-demo the way he is anti-benchmax; he still wanted the sore-wrist Whisper Flow version if it can be made reliable. Games, auto-battlers, Stardew Valley NPCs: he is a fan and probably too busy, and Jev in the game loop is probably too expensive, but state machines for NPCs are sitting there.

Claude Code and Codex, as he roughly ranked them, are built around a one-model world. Open coding agents are “getting their Jevon” and still roughly at par, because there is only so much you can do with a while-loop. The first killer use case that requires this architecture becomes a temporary monopoly the open agents can copy. He does not know what the single-model products will do. Integrating with them would be lovely. Competing with them is not his job as infrastructure. He has an internal design-patterns doc he hoped to share after walking home, if the team does not veto him. He is, he noted, not in charge.

He would not pre-train with a billion dollars

Mid-training fascinates him as a spectrum and a cost-saving alternative to pre-training again. Rapid fine-tuning sits on the surface; repetition bakes capabilities deeper until they are robust; System One is what gets robust. He is a fan of all forms of training and would not do them all himself, because they are expensive. Anything he would say to an investor he wants said in public: if you gave him a billion dollars, he still would not pre-train. An AI engineer can slice, dice, and Frankenstein. It is not elegant. It solves problems.

He hates fracturing intelligence in post-training. Chat-first reasoning mode forces it. RLHF’s natural complaints — sycophancy, overconfidence, hallucination, LMSYS-style bold-italic-emoji house style, the long write-up and a follow-up question so it feels human — come because strings are weird and the reward model punishes visible derailment so hard that the model must be miscalibrated and mode-dropped to stay on the rails. That warped distribution then interacts with reasoning models, which he described as simple linear things that tend to cheat. Two objectives is already the act of fracturing. He will not bake “you are Jev from TypeSafe” into the weights. Identity is a first-party-product problem. An API should represent what the internet thinks and be correct. If people jailbreak it into saying it is OpenAI or Qwen or Claude, that is closer to truth than a corporate mask.

Omni-models versus broken-out capabilities: empirical. Sometimes other modalities help, sometimes they hurt; labs seem to be moving away from speech as distinct from audio because it does not generalize. Computer use is not currently solved; there might be no amount of data that solves it without better methods. Scaling laws are not “throw money.” They are “how good is the thing,” and some worlds never get good enough.

InstructGPT took half the market and wrote slop

The problem was in his head before ChatGPT launched. The ChatGPT team, he said, was doing the thing researchers are bad at and successful product people are good at: caring about the experience. He fought to deploy InstructGPT. Early versions used an unpublished algorithm he made because cleaning PPO data was too slow. It immediately took 50% of the LLM market at the time. He went through the launch video to make sure that history was true. It looked so much like AGI — superhuman instruction in, instruction out — that everyone should have an answer for why it was not. His answer, after the fact: it got used for Jasper and Copy.ai, slop on web pages. They worried they had made the internet worse. He went back to the drawing board. Work backwards from an AI-based economic revolution. When it happens, if AI is an API, who is calling it — humans or code? Many nines of code. All the optimization was going into the humans. That was the north star. He wrote a document, talked to Sam Altman, and Sam said it was so good he should go work on it. Almeida said he had a job. He assumed Anthropic must already be doing it and OpenAI would be cooked; OpenAI, in his telling, is better at catching up than innovating. ChatGPT copied Claude. Claude did coding agents. The instruction-following team eventually declared victory. He started training models. He thought it would take a week. He was “unbelievably wrong,” and he apologized, on the record, to everyone at OpenAI who had to live next to that confidence.

Around Thanksgiving, during the coup — “safety took over the company,” tea for another time — he canceled everything, took OpenAI GPUs while people were on holiday, and ran. Signs of life, not deployable. If an AI winter happened and he had not gone all-in, he would consider himself personally responsible both for RLHF’s widening of overpromise versus under-delivery and for not pursuing this. Other companies said a startup would be faster than starting a lab inside. He called Eric first. He did not try to recruit Sasha; he asked if he was crazy, if the OpenAI bubble had hidden a solution. She said she was in, folded her startup after he told her to think about it, and joined. Within two weeks they had funding and people living in his apartment. He is a neat freak. It was the worst.

Most neo-labs, he said, are crap. He does not want them as peers. He does not particularly value researchers except insofar as they care about the right task. Pure research pedigree generally does not create value. Neo-labs redo work from scratch with a low probability of moving the frontier; the ones he has talked to often want money to play with experiments and have no direction. If they have a direction, he is in favor. If you want to play with research, a big lab is probably the right place. If you want to break the unimodal mind, please do. He would rather stay a pure technologist than get good at misleading the public “for their own good,” a pattern he thinks COVID-era appeals to authority made worse. swyx joked that Jev should run for president; Almeida said he already trusts Jev’s decisions over his own. He described himself as “zeroth-percent entrepreneurial.” He never wanted to be a CEO. He cannot imagine anyone doing it twice. An investor asked which CEOs he looks up to. He said, “Ew.” At OpenAI he felt disempowered in an insane house where everything was how to put ChatGPT on something. Function calling, in his view, is an anti-developer interface: a hack on a hack. He used to ask to be removed from any function-calling project that did not expose a per-function logit — a probability, a confidence. Disney and AI Dungeon need different refusal thresholds. The only control in function calling is “pretty please” in a system message. Skills have not fixed that. Coding agents overfit their harnesses and beg tools to be called. That is the interface he is trying to make obsolete.

The research he wants other people to do, because he will be busy for fifty years: intelligent games; and coding agents freed from the tyranny of the KV cache. He wrote “KV cache rules everything around me” because, he claimed, nobody had used that phrase with that spelling. The cache locks you into one model and into append-only context, which is the enemy of state management, abstraction, and decomposition. Sub-agents fail because the state you would need to pass is more expensive to read than the task is worth. If context lookup were cheap, you could search a labeled tree of subtasks, read historical context instead of reinventing continual learning as a memory-management problem, let parallel agents share computer memory with real locks, and have read-only observers summarize without re-eating the transcript. Jev to solve locks, swyx joked. Almeida did not treat it as a joke. He pointed, with swyx, at work like Prime Intellect’s agent / recursive language models as early and not yet popular. He wants to fund that kind of play once they figure out credits.

They are hiring data people they refuse to treat as a slur, platform people because speed of light is now the bottleneck — he is sad European users get 3× instead of 100× — and people to ship more shapes of intelligence than Jev. The goal is not a one-trick pony. He wants an AWS of intelligence, System One as TCP, “seven more layers to go.” swyx had already pitched Temporal as layer eight. Almeida had to go back to work, or sleep. He said it was his first time on the show, and that the launch had already changed the path of technological history. If TypeSafe disappeared, he guessed it might take a year or two for others to catch up, unless model quality matters, in which case they will be in a good position for a long time. He does not actually know. Both sentences were allowed to stand.