← Back
Theo September 21, 2026 30m

Jev is incredible

Read full transcript 24 segments
  1. Oh boy, it's new model time. This one's Oh boy, it's new model time. This one's very different though. This isn't our very different though. This isn't our very different though. This isn't our usual new LLM that's slightly better at usual new LLM that's slightly better at usual new LLM that's slightly better at code and it's not going to be the type code and it's not going to be the type code and it's not going to be the type of thing that the model counter guy is of thing that the model counter guy is of thing that the model counter guy is going to show up and say, "Look, best going to show up and say, "Look, best going to show up and say, "Look, best new model." If anything, this might take new model." If anything, this might take new model." If anything, this might take the counter guy out of his job because the counter guy out of his job because the counter guy out of his job because this model is for data processing. It's this model is for data processing. It's this model is for data processing. It's by a company called Typesafe AI and it's by a company called Typesafe AI and it's by a company called Typesafe AI and it's called Jev. The point of this model called Jev. The point of this model called Jev. The point of this model isn't to generate text, write code, or isn't to generate text, write code, or isn't to generate text, write code, or do all the things that we expect models do all the things that we expect models do all the things that we expect models to do today. It is classification. This to do today. It is classification. This to do today. It is classification. This model is the best model ever made to model is the best model ever made to model is the best model ever made to take data and organize it and classify take data and organize it and classify take data and organize it and classify it. It returns JSON perfectly in a it. It returns JSON perfectly in a it. It returns JSON perfectly in a type-S safe format. When you give it a type-S safe format. When you give it a type-S safe format. When you give it a format, it does it and it does it well. format, it does it and it does it well. format, it does it and it does it well. That means this model isn't going to That means this model isn't going to That means this model isn't going to replace things like Fable or Astra. replace things like Fable or Astra. replace things like Fable or Astra. Well, hopefully not. If you are using Well, hopefully not. If you are using Well, hopefully not. If you are using Fable and Astra for the stuff you can do Fable and Astra for the stuff you can do Fable and Astra for the stuff you can do with this, I have a lot of questions for with this, I have a lot of questions for with this, I have a lot of questions for you. All that said, this model is you. All that said, this model is you. All that said, this model is incredibly capable, and that's why it's incredibly capable, and that's why it's incredibly capable, and that's why it's blown up on Twitter. I've seen demos of blown up on Twitter. I've seen demos of blown up on Twitter. I've seen demos of everything from lightning fast computer everything from lightning fast computer everything from lightning fast computer use even in like the iOS simulator to use even in like the iOS simulator to use even in like the iOS simulator to the model playing Minecraft at lightning the model playing Minecraft at lightning the model playing Minecraft at lightning speed to compaction that takes under a speed to compaction that takes under a speed to compaction that takes under a second instead of multiple minutes. I second instead of multiple minutes. I second instead of multiple minutes. I can't wait to show you all the best can't wait to show you all the best can't wait to show you all the best things you can do with this model. But things you can do with this model. But things you can do with this model. But first, we got to classify something. The first, we got to classify something. The first, we got to classify something. The next section, which is the sponsor next section, which is the sponsor next section, which is the sponsor break. If you're not a developer, you break. If you're not a developer, you break. If you're not a developer, you can skip this ad. But if you are, you can skip this ad. But if you are, you can skip this ad. But if you are, you should listen close. Have you ever sent should listen close. Have you ever sent should listen close. Have you ever sent a prompt to Cloud Code, Codeex, Cursor, a prompt to Cloud Code, Codeex, Cursor, a prompt to Cloud Code, Codeex, Cursor, or some other tool and been surprised or some other tool and been surprised or some other tool and been surprised that it took hours when you expected it that it took hours when you expected it that it took hours when you expected it to take minutes? Did you check to see to take minutes? Did you check to see to take minutes? Did you check to see why? Because I would bet, good money, why? Because I would bet, good money, why? Because I would bet, good money, there's a very good chance that a there's a very good chance that a there's a very good chance that a handful of things happened. Maybe it handful of things happened. Maybe it handful of things happened. Maybe it took too long to download the Docker took too long to download the Docker took too long to download the Docker image. Maybe it spun up and had some image. Maybe it spun up and had some image. Maybe it spun up and had some error when it did. Maybe the CI that it error when it did. Maybe the CI that it error when it did. Maybe the CI that it was trying to run took forever. Maybe it was trying to run took forever. Maybe it was trying to run took forever. Maybe it pushed it up to GitHub and the CI run on pushed it up to GitHub and the CI run on pushed it up to GitHub and the CI run on there was going to take hours long. I

  2. there was going to take hours long. I there was going to take hours long. I cannot tell you how many times I've run cannot tell you how many times I've run cannot tell you how many times I've run into this. Wouldn't it be nice if your into this. Wouldn't it be nice if your into this. Wouldn't it be nice if your agents could do all of those things way agents could do all of those things way agents could do all of those things way faster and more reliably, and when those faster and more reliably, and when those faster and more reliably, and when those things failed, you got insight into it? things failed, you got insight into it? things failed, you got insight into it? Imagine taking all of these 8-minute Imagine taking all of these 8-minute Imagine taking all of these 8-minute builds and knocking them down to 20 builds and knocking them down to 20 builds and knocking them down to 20 seconds or building and testing three seconds or building and testing three seconds or building and testing three React apps in seven Go binaries in 30 React apps in seven Go binaries in 30 React apps in seven Go binaries in 30 seconds or making your crossplatform seconds or making your crossplatform seconds or making your crossplatform builds 10 to 20 times faster. Hopefully, builds 10 to 20 times faster. Hopefully, builds 10 to 20 times faster. Hopefully, you get the idea now because today's you get the idea now because today's you get the idea now because today's sponsor is depot and they are trying sponsor is depot and they are trying sponsor is depot and they are trying their damnedest to make AI more capable, their damnedest to make AI more capable, their damnedest to make AI more capable, not by giving it things it can't do not by giving it things it can't do not by giving it things it can't do already, but by making all of the things already, but by making all of the things already, but by making all of the things it does do more reliable and faster. it does do more reliable and faster. it does do more reliable and faster. Whether you're looking for a faster and Whether you're looking for a faster and Whether you're looking for a faster and cheaper alternative to GitHub actions or cheaper alternative to GitHub actions or cheaper alternative to GitHub actions or just a nice place to let your agents run just a nice place to let your agents run just a nice place to let your agents run or a cache for all the Docker images you or a cache for all the Docker images you or a cache for all the Docker images you and your team are using, Dep Depot and your team are using, Dep Depot and your team are using, Dep Depot provides all of this and more. Their CI provides all of this and more. Their CI provides all of this and more. Their CI platform is fully compatible with GitHub platform is fully compatible with GitHub platform is fully compatible with GitHub actions, which means you can swap over actions, which means you can swap over actions, which means you can swap over trivially. But if you don't want to deal trivially. But if you don't want to deal trivially. But if you don't want to deal with GitHub's downtime, you can move with GitHub's downtime, you can move with GitHub's downtime, you can move fully over to depot's platform. Still fully over to depot's platform. Still fully over to depot's platform. Still compatible with actions, but without compatible with actions, but without compatible with actions, but without having to worry about GitHub's having to worry about GitHub's having to worry about GitHub's reliability. You can also transform reliability. You can also transform reliability. You can also transform those workflows to let them run in those workflows to let them run in those workflows to let them run in parallel, which ends up being comically parallel, which ends up being comically parallel, which ends up being comically faster than the alternatives on GitHub.

  3. faster than the alternatives on GitHub. faster than the alternatives on GitHub. Their CLI is incredible, too. You and Their CLI is incredible, too. You and Their CLI is incredible, too. You and your agents will understand it your agents will understand it your agents will understand it immediately. It makes it trivial to immediately. It makes it trivial to immediately. It makes it trivial to migrate your CI, run it without pushing migrate your CI, run it without pushing migrate your CI, run it without pushing up code, diagnose things when they go up code, diagnose things when they go up code, diagnose things when they go wrong, find secrets and clone them wrong, find secrets and clone them wrong, find secrets and clone them across different places, and more. You across different places, and more. You across different places, and more. You can even SSH into a session while it is can even SSH into a session while it is can even SSH into a session while it is running to figure out what's going wrong running to figure out what's going wrong running to figure out what's going wrong in real time. And when I say you, in real time. And when I say you, in real time. And when I say you, obviously, I mean your agents. GitHub obviously, I mean your agents. GitHub obviously, I mean your agents. GitHub actions were poorly assembled for humans actions were poorly assembled for humans actions were poorly assembled for humans over 10 years ago. Depot was crafted over 10 years ago. Depot was crafted over 10 years ago. Depot was crafted carefully for agents today. Figure out carefully for agents today. Figure out carefully for agents today. Figure out what that means at soyv.link/depo. what that means at soyv.link/depo. what that means at soyv.link/depo. Quick breakdown on how this all happened Quick breakdown on how this all happened Quick breakdown on how this all happened because it's actually been pretty crazy because it's actually been pretty crazy because it's actually been pretty crazy to watch. It feels like it just poof to watch. It feels like it just poof to watch. It feels like it just poof appeared and took over my entire Twitter appeared and took over my entire Twitter appeared and took over my entire Twitter feed. Like half of what I've been seeing feed. Like half of what I've been seeing feed. Like half of what I've been seeing is all of the stuff about Jev. It was is all of the stuff about Jev. It was is all of the stuff about Jev. It was created by Dio who used to work at created by Dio who used to work at created by Dio who used to work at OpenAI, he apparently helped co-invent OpenAI, he apparently helped co-invent OpenAI, he apparently helped co-invent chat GPT as well as RHF. And as great as chat GPT as well as RHF. And as great as chat GPT as well as RHF. And as great as RHF is for model behaviors, it is not RHF is for model behaviors, it is not RHF is for model behaviors, it is not helping as much with classification, helping as much with classification, helping as much with classification, which is the thing he cares about. If which is the thing he cares about. If which is the thing he cares about. If you focus more on the classifying side you focus more on the classifying side you focus more on the classifying side and less on the general usage of text and less on the general usage of text and less on the general usage of text generation side, what you can get is way generation side, what you can get is way generation side, what you can get is way way faster and way way cheaper. In way faster and way way cheaper. In way faster and way way cheaper. In particular, all output tokens being particular, all output tokens being particular, all output tokens being free. The cost characteristics of this free. The cost characteristics of this free. The cost characteristics of this model are almost as insane as the speed.

  4. model are almost as insane as the speed. model are almost as insane as the speed. Both are crazy. So, let's read through Both are crazy. So, let's read through Both are crazy. So, let's read through the announcement. The core concept here the announcement. The core concept here the announcement. The core concept here is system one models. This is an idea is system one models. This is an idea is system one models. This is an idea that comes from thinking fast and slow that comes from thinking fast and slow that comes from thinking fast and slow by Daniel Conman. System one is the fast by Daniel Conman. System one is the fast by Daniel Conman. System one is the fast and intuitive judgment of your brain. and intuitive judgment of your brain. and intuitive judgment of your brain. the part of your brain that can do the part of your brain that can do the part of your brain that can do things without effort that looks at things without effort that looks at things without effort that looks at somebody's like, "Yeah, that shirt's somebody's like, "Yeah, that shirt's somebody's like, "Yeah, that shirt's blue." System two is slow and deliberate blue." System two is slow and deliberate blue." System two is slow and deliberate reasoning. In type safe's framing, Jev reasoning. In type safe's framing, Jev reasoning. In type safe's framing, Jev handles the first kind of task while handles the first kind of task while handles the first kind of task while reasoning models still handle the reasoning models still handle the reasoning models still handle the second. They named it Jeb after Jevan's second. They named it Jeb after Jevan's second. They named it Jeb after Jevan's paradox, which I actually think is paradox, which I actually think is paradox, which I actually think is pretty cute. The idea is that when the pretty cute. The idea is that when the pretty cute. The idea is that when the steam engine gets more efficient, it steam engine gets more efficient, it steam engine gets more efficient, it actually increased coal consumption, not actually increased coal consumption, not actually increased coal consumption, not decreased, because now that power is decreased, because now that power is decreased, because now that power is cheaper, we can use it for more things. cheaper, we can use it for more things. cheaper, we can use it for more things. That is really the right way to think of That is really the right way to think of That is really the right way to think of this model. It's not unlocking new this model. It's not unlocking new this model. It's not unlocking new capabilities that weren't possible capabilities that weren't possible capabilities that weren't possible before. It's making them fast and cheap before. It's making them fast and cheap before. It's making them fast and cheap enough that they're way more viable than enough that they're way more viable than enough that they're way more viable than they've been. But what are those cases? they've been. But what are those cases? they've been. But what are those cases? What is this valuable for? Pretty much What is this valuable for? Pretty much What is this valuable for? Pretty much anything that needs to be classified. anything that needs to be classified. anything that needs to be classified. Think organizing videos by their topic Think organizing videos by their topic Think organizing videos by their topic in my YouTube channel. Figuring out if in my YouTube channel. Figuring out if in my YouTube channel. Figuring out if an email is important or not and what an email is important or not and what an email is important or not and what category it should be put under. Things category it should be put under. Things category it should be put under. Things like safety and moderation, identifying like safety and moderation, identifying like safety and moderation, identifying what messages are safe and which ones what messages are safe and which ones what messages are safe and which ones aren't and what makes them unsafe.

  5. aren't and what makes them unsafe. aren't and what makes them unsafe. There's so many things you can use it There's so many things you can use it There's so many things you can use it for. And when it's as insanely fast as for. And when it's as insanely fast as for. And when it's as insanely fast as this model is, the results are kind of this model is, the results are kind of this model is, the results are kind of crazy. We'll go to the official crazy. We'll go to the official crazy. We'll go to the official announcement in a sec, but first I just announcement in a sec, but first I just announcement in a sec, but first I just want to show you how insanely fast the want to show you how insanely fast the want to show you how insanely fast the model is with an example of something it model is with an example of something it model is with an example of something it can do. Not saying it can do it well. can do. Not saying it can do it well. can do. Not saying it can do it well. You'll actually see the type 1, type two You'll actually see the type 1, type two You'll actually see the type 1, type two distinction here. But this is a game of distinction here. But this is a game of distinction here. But this is a game of checkers that I built to use Jev. Jev checkers that I built to use Jev. Jev checkers that I built to use Jev. Jev takes the whole board as an input, as a takes the whole board as an input, as a takes the whole board as an input, as a state, not an image because it doesn't state, not an image because it doesn't state, not an image because it doesn't have vision. It takes the board state as have vision. It takes the board state as have vision. It takes the board state as data and then it chooses what to move data and then it chooses what to move data and then it chooses what to move based on the output it gives. So here it based on the output it gives. So here it based on the output it gives. So here it might say it wants to move B6 to C5. I might say it wants to move B6 to C5. I might say it wants to move B6 to C5. I go first, so I'll make a move. Watch how go first, so I'll make a move. Watch how go first, so I'll make a move. Watch how fast it responds. I'm clicking now. fast it responds. I'm clicking now. fast it responds. I'm clicking now. Yeah, it's practically instant. There is a catch though. It's not very There is a catch though. It's not very smart. smart. smart. I was able to barely pay attention while I was able to barely pay attention while I was able to barely pay attention while playing and crush this model. Yeah, it's playing and crush this model. Yeah, it's playing and crush this model. Yeah, it's just it's not good. I was barely paying just it's not good. I was barely paying just it's not good. I was barely paying attention and I'm going to crush it attention and I'm going to crush it attention and I'm going to crush it here. Point I'm trying to make is that here. Point I'm trying to make is that here. Point I'm trying to make is that it's super fast but it's not using the it's super fast but it's not using the it's super fast but it's not using the reasoning part of its brain so to speak.

  6. reasoning part of its brain so to speak. reasoning part of its brain so to speak. It is just processing the data. Let's It is just processing the data. Let's It is just processing the data. Let's take a look at what do models have been take a look at what do models have been take a look at what do models have been superhuman at chat for years. So where's superhuman at chat for years. So where's superhuman at chat for years. So where's all the automation? This has been my all the automation? This has been my all the automation? This has been my driving question for the last four driving question for the last four driving question for the last four years. At OpenAI, I help build the years. At OpenAI, I help build the years. At OpenAI, I help build the methods that make language models useful methods that make language models useful methods that make language models useful at following instructions and talking at following instructions and talking at following instructions and talking with people. The work ended up as the with people. The work ended up as the with people. The work ended up as the research behind Chaz GBT. At the time, I research behind Chaz GBT. At the time, I research behind Chaz GBT. At the time, I thought maybe chat models would lead to thought maybe chat models would lead to thought maybe chat models would lead to AGI. But despite the hype, it became AGI. But despite the hype, it became AGI. But despite the hype, it became obvious to me that there was something obvious to me that there was something obvious to me that there was something really big missing. After two years in really big missing. After two years in really big missing. After two years in stealth, countless technical challenges, stealth, countless technical challenges, stealth, countless technical challenges, and research breakthroughs, he's beyond and research breakthroughs, he's beyond and research breakthroughs, he's beyond excited to announce that today Typesafe excited to announce that today Typesafe excited to announce that today Typesafe AI is releasing its first system 1 AI is releasing its first system 1 AI is releasing its first system 1 model, a new class of frontier model model, a new class of frontier model model, a new class of frontier model built to make fast structured decisions built to make fast structured decisions built to make fast structured decisions that software can use directly. This is that software can use directly. This is that software can use directly. This is one of the biggest important pieces to one of the biggest important pieces to one of the biggest important pieces to understand. The point of this model is understand. The point of this model is understand. The point of this model is to integrate it into your code. This to integrate it into your code. This to integrate it into your code. This model isn't even useful if a human model isn't even useful if a human model isn't even useful if a human triggers it. The model is useful if the triggers it. The model is useful if the triggers it. The model is useful if the human triggers some code to run or human triggers some code to run or human triggers some code to run or something else triggers some code to run something else triggers some code to run something else triggers some code to run and when the code runs it passes certain and when the code runs it passes certain and when the code runs it passes certain data to this model and then it comes data to this model and then it comes data to this model and then it comes back with JSON. That's when it's useful back with JSON. That's when it's useful back with JSON. That's when it's useful when it's integrated almost like a when it's integrated almost like a when it's integrated almost like a function in your codebase. In order to function in your codebase. In order to function in your codebase. In order to do this they had to build a new stack do this they had to build a new stack do this they had to build a new stack entirely focused on automation, new entirely focused on automation, new entirely focused on automation, new model architectures, parallel samplers model architectures, parallel samplers model architectures, parallel samplers for maximum efficiency and training for maximum efficiency and training for maximum efficiency and training methods that they call reinforcement methods that they call reinforcement methods that they call reinforcement learning for calibrated decisions. The learning for calibrated decisions. The learning for calibrated decisions. The first public version of this is Jev first public version of this is Jev first public version of this is Jev which is now available in early access.

  7. which is now available in early access. which is now available in early access. It is invite only, but you can get It is invite only, but you can get It is invite only, but you can get access to it on things like open route access to it on things like open route access to it on things like open route or the versel AI gateway and a couple or the versel AI gateway and a couple or the versel AI gateway and a couple other places. While Jev gives up string other places. While Jev gives up string other places. While Jev gives up string generation, it's optimized for generation, it's optimized for generation, it's optimized for structural outputs and it can't structural outputs and it can't structural outputs and it can't hallucinate. This is a key piece. Since hallucinate. This is a key piece. Since hallucinate. This is a key piece. Since it has to give the output in a certain it has to give the output in a certain it has to give the output in a certain format, it's not going to change or make format, it's not going to change or make format, it's not going to change or make up the format, which happens a hilarious up the format, which happens a hilarious up the format, which happens a hilarious amount. Think of Jev as a frontier amount. Think of Jev as a frontier amount. Think of Jev as a frontier intelligence function call. Unstructured intelligence function call. Unstructured intelligence function call. Unstructured state in, typed probabilistic decisions state in, typed probabilistic decisions state in, typed probabilistic decisions out. That's the key piece. You hand it out. That's the key piece. You hand it out. That's the key piece. You hand it some text data and you get back some text data and you get back some text data and you get back a type- safe shape. There have been a type- safe shape. There have been a type- safe shape. There have been other attempts to do this in various other attempts to do this in various other attempts to do this in various different ways from attempts to train different ways from attempts to train different ways from attempts to train models to act like this to tools to models to act like this to tools to models to act like this to tools to force other models to behave similarly force other models to behave similarly force other models to behave similarly too. One that I really liked once it too. One that I really liked once it too. One that I really liked once it clicked for me is called BAML. The point clicked for me is called BAML. The point clicked for me is called BAML. The point of BAML was to make a language that of BAML was to make a language that of BAML was to make a language that interfaces between agents and real code interfaces between agents and real code interfaces between agents and real code so that you can define something in a so that you can define something in a so that you can define something in a syntax that of course agents understand syntax that of course agents understand syntax that of course agents understand but also to give a specific format and but also to give a specific format and but also to give a specific format and instruction set to the model to get it instruction set to the model to get it instruction set to the model to get it to output a certain shape. They frame it to output a certain shape. They frame it to output a certain shape. They frame it kind of like TypeScript but as the kind of like TypeScript but as the kind of like TypeScript but as the interface from the TypeScript to the interface from the TypeScript to the interface from the TypeScript to the LLM. For example, here is a text LLM. For example, here is a text LLM. For example, here is a text sentiment classifier that they have on sentiment classifier that they have on sentiment classifier that they have on their site. You give it a label type their site. You give it a label type their site. You give it a label type which positive, negative or neutral are which positive, negative or neutral are which positive, negative or neutral are valid for verdict which has label and valid for verdict which has label and valid for verdict which has label and confidence which is a float. Then you confidence which is a float. Then you confidence which is a float. Then you define the function classify. It takes define the function classify. It takes define the function classify. It takes in text. It outputs a verdict. You use in text. It outputs a verdict. You use in text. It outputs a verdict. You use OpenAI GD55 here as the client. You hand OpenAI GD55 here as the client. You hand OpenAI GD55 here as the client. You hand it the prompt. You use their special it the prompt. You use their special it the prompt. You use their special formatting things. And then you can call formatting things. And then you can call formatting things. And then you can call this in your TypeScript code in order to this in your TypeScript code in order to this in your TypeScript code in order to get a formatted correctly output. BAML get a formatted correctly output. BAML get a formatted correctly output. BAML is basically a madeup language, but what

  8. is basically a madeup language, but what is basically a madeup language, but what makes it cool is how well it interfaces makes it cool is how well it interfaces makes it cool is how well it interfaces with other languages and how it fixes with other languages and how it fixes with other languages and how it fixes all the things that can go wrong. For all the things that can go wrong. For all the things that can go wrong. For example, I was using this a bunch with example, I was using this a bunch with example, I was using this a bunch with GBT OSS120B to find comments that GBT OSS120B to find comments that GBT OSS120B to find comments that mentioned sponsors and things. And when mentioned sponsors and things. And when mentioned sponsors and things. And when I did that, I had a really good I did that, I had a really good I did that, I had a really good experience with it. It all just worked. experience with it. It all just worked. experience with it. It all just worked. And when I tried to move off BAML for And when I tried to move off BAML for And when I tried to move off BAML for another thing I was doing, I learned another thing I was doing, I learned another thing I was doing, I learned that apparently GPOSS120B that apparently GPOSS120B that apparently GPOSS120B doesn't format JSON correctly half the doesn't format JSON correctly half the doesn't format JSON correctly half the time. One of the many things they do time. One of the many things they do time. One of the many things they do with BAML is in their runtime, they fix with BAML is in their runtime, they fix with BAML is in their runtime, they fix the mal formatting in the JSON outputs the mal formatting in the JSON outputs the mal formatting in the JSON outputs to make sure it is guaranteed to match to make sure it is guaranteed to match to make sure it is guaranteed to match the format it puts on the tin. when you the format it puts on the tin. when you the format it puts on the tin. when you say this is the response format and you say this is the response format and you say this is the response format and you call it through BAML that's the format call it through BAML that's the format call it through BAML that's the format you get sometimes that takes time though you get sometimes that takes time though you get sometimes that takes time though because it has to reformat it fix it because it has to reformat it fix it because it has to reformat it fix it change it all with code that it's change it all with code that it's change it all with code that it's running but and even some high-end running but and even some high-end running but and even some high-end models can sometimes screw these things models can sometimes screw these things models can sometimes screw these things up literally can't if you give it a JSON up literally can't if you give it a JSON up literally can't if you give it a JSON format you get back that JSON format it format you get back that JSON format it format you get back that JSON format it doesn't handle things like a message doesn't handle things like a message doesn't handle things like a message history well as input it's really history well as input it's really history well as input it's really looking for structured program state looking for structured program state looking for structured program state like data it can use to make decisions I like data it can use to make decisions I like data it can use to make decisions I already see people getting confused so I already see people getting confused so I already see people getting confused so I want to be really really clear here the want to be really really clear here the want to be really really clear here the point of things like BAML, like Jev, and point of things like BAML, like Jev, and point of things like BAML, like Jev, and like all structured outputs is because like all structured outputs is because like all structured outputs is because LLMs are not deterministic. If I have LLMs are not deterministic. If I have LLMs are not deterministic. If I have code that formats someone's name, it code that formats someone's name, it code that formats someone's name, it will always do the same thing when I will always do the same thing when I will always do the same thing when I call it with a given name. If I have an call it with a given name. If I have an call it with a given name. If I have an LLM that I ask to format the name, it LLM that I ask to format the name, it LLM that I ask to format the name, it might do something else. If I have a might do something else. If I have a might do something else. If I have a function that returns a user object that function that returns a user object that function that returns a user object that has a name colon, string, and an age has a name colon, string, and an age has a name colon, string, and an age colon, number, and I write code to get colon, number, and I write code to get colon, number, and I write code to get that, it will always be that shape. If I that, it will always be that shape. If I that, it will always be that shape. If I call an LLM, it might screw up the

  9. call an LLM, it might screw up the call an LLM, it might screw up the formatting of the name. It might use a formatting of the name. It might use a formatting of the name. It might use a float for age instead of an int. There's float for age instead of an int. There's float for age instead of an int. There's a lot of different things that can go a lot of different things that can go a lot of different things that can go wrong. The point of Jev is that you can wrong. The point of Jev is that you can wrong. The point of Jev is that you can call it as reliably as you call code and call it as reliably as you call code and call it as reliably as you call code and it's really fast, kind of like code. It it's really fast, kind of like code. It it's really fast, kind of like code. It also can do this for a bunch of inputs also can do this for a bunch of inputs also can do this for a bunch of inputs at once in parallel, which is super at once in parallel, which is super at once in parallel, which is super super cool. And the token cost is super cool. And the token cost is super cool. And the token cost is insanely low. It's about 4 cents per insanely low. It's about 4 cents per insanely low. It's about 4 cents per million tokens in compared to $10 per million tokens in compared to $10 per million tokens in compared to $10 per million on a model like Fable. I see million on a model like Fable. I see million on a model like Fable. I see more confusion in chat already. So, is more confusion in chat already. So, is more confusion in chat already. So, is it deterministic? The format of the it deterministic? The format of the it deterministic? The format of the output is yes. There are two ways things output is yes. There are two ways things output is yes. There are two ways things can be deterministic. The shape and the can be deterministic. The shape and the can be deterministic. The shape and the content. The content isn't content. The content isn't content. The content isn't deterministic, but it should be quite deterministic, but it should be quite deterministic, but it should be quite reliable. The actual format is what is reliable. The actual format is what is reliable. The actual format is what is deterministic. It will always give you deterministic. It will always give you deterministic. It will always give you the format that you define. And on the the format that you define. And on the the format that you define. And on the topic of cost, the tokens out are free topic of cost, the tokens out are free topic of cost, the tokens out are free because it's so cheap. They frame it as because it's so cheap. They frame it as because it's so cheap. They frame it as too cheap to meter. Meanwhile, with too cheap to meter. Meanwhile, with too cheap to meter. Meanwhile, with LLMs, the output tokens are so LLMs, the output tokens are so LLMs, the output tokens are so expensive. It's often the majority cost expensive. It's often the majority cost expensive. It's often the majority cost when you're doing stuff like this. when you're doing stuff like this. when you're doing stuff like this. Traditional LMS can take three to over Traditional LMS can take three to over Traditional LMS can take three to over 300 seconds to do the type of 300 seconds to do the type of 300 seconds to do the type of classification work that this model can classification work that this model can classification work that this model can do in 70 to 500 milliseconds. That is do in 70 to 500 milliseconds. That is do in 70 to 500 milliseconds. That is actually 40 to 200x faster. There's no actually 40 to 200x faster. There's no actually 40 to 200x faster. There's no doubt on that claim. That is real and doubt on that claim. That is real and doubt on that claim. That is real and true and based. Even if prompted for a true and based. Even if prompted for a true and based. Even if prompted for a confidence estimate, models tend to be confidence estimate, models tend to be confidence estimate, models tend to be overconfident and inconsistent. The overconfident and inconsistent. The overconfident and inconsistent. The model can do a task 95% of the time, but model can do a task 95% of the time, but model can do a task 95% of the time, but doesn't say when it's in the 5%, you doesn't say when it's in the 5%, you doesn't say when it's in the 5%, you can't automate the task. Meanwhile, Jev can't automate the task. Meanwhile, Jev can't automate the task. Meanwhile, Jev will always communicate confidence and will always communicate confidence and will always communicate confidence and uncertainty with every output. So you uncertainty with every output. So you uncertainty with every output. So you can calibrate around the accuracy can calibrate around the accuracy can calibrate around the accuracy numbers you get and the answers also numbers you get and the answers also numbers you get and the answers also tend to be more consistent. I like the tend to be more consistent. I like the tend to be more consistent. I like the framing here of one of the good use

  10. framing here of one of the good use framing here of one of the good use cases being a smart if statement. I know cases being a smart if statement. I know cases being a smart if statement. I know a lot of y'all don't code anymore sadly. a lot of y'all don't code anymore sadly. a lot of y'all don't code anymore sadly. I get it. But if statements were a great I get it. But if statements were a great I get it. But if statements were a great way of thinking of logic. And that's way of thinking of logic. And that's way of thinking of logic. And that's what this model is for as a thing you what this model is for as a thing you what this model is for as a thing you put between states to decide what path put between states to decide what path put between states to decide what path you go down or to map reduce over a you go down or to map reduce over a you go down or to map reduce over a giant pile of data in order to get giant pile of data in order to get giant pile of data in order to get insights out of it. And this is one of insights out of it. And this is one of insights out of it. And this is one of the coolest things you can do with it. the coolest things you can do with it. the coolest things you can do with it. real-time applications. Not that like real-time applications. Not that like real-time applications. Not that like it'll build the app, but is a tool in it'll build the app, but is a tool in it'll build the app, but is a tool in the app because it responds in just the app because it responds in just the app because it responds in just under 500 milliseconds worst case. You under 500 milliseconds worst case. You under 500 milliseconds worst case. You can do something silly like a search can do something silly like a search can do something silly like a search with it or the demo I just showed with with it or the demo I just showed with with it or the demo I just showed with chess or checkers. This demo from Matt chess or checkers. This demo from Matt chess or checkers. This demo from Matt is really cool as well where you can is really cool as well where you can is really cool as well where you can define a term and ask it to generate a define a term and ask it to generate a define a term and ask it to generate a color palette effectively for it and it color palette effectively for it and it color palette effectively for it and it shifts that bar at the bottom and it's shifts that bar at the bottom and it's shifts that bar at the bottom and it's practically real time because the model practically real time because the model practically real time because the model is so quick to respond. The final use is so quick to respond. The final use is so quick to respond. The final use case they have in their example list case they have in their example list case they have in their example list here is verifying everything. Score, here is verifying everything. Score, here is verifying everything. Score, judge, verify, guard rail, and detect judge, verify, guard rail, and detect judge, verify, guard rail, and detect jailbreaks of LLM prompts, reasoning jailbreaks of LLM prompts, reasoning jailbreaks of LLM prompts, reasoning traces, and outputs. Stuff like that. I traces, and outputs. Stuff like that. I traces, and outputs. Stuff like that. I want to make a quick point though want to make a quick point though want to make a quick point though because one of the questions I've seen because one of the questions I've seen because one of the questions I've seen the most by far is something along the the most by far is something along the the most by far is something along the lines of, "Wait, so if I use four lines of, "Wait, so if I use four lines of, "Wait, so if I use four different LLMs to generate an output, I different LLMs to generate an output, I different LLMs to generate an output, I could use this to judge it." you do could use this to judge it." you do could use this to judge it." you do understand how expensive it would be to understand how expensive it would be to understand how expensive it would be to generate all four of those and how silly generate all four of those and how silly generate all four of those and how silly it would be to have a model that's only it would be to have a model that's only it would be to have a model that's only using one side of its brain judge that using one side of its brain judge that using one side of its brain judge that the benefit of other LLMs is that they the benefit of other LLMs is that they the benefit of other LLMs is that they can reason. They can think through the can reason. They can think through the can reason. They can think through the decision. They can walk through your decision. They can walk through your decision. They can walk through your codebase using agentic tools in order to codebase using agentic tools in order to codebase using agentic tools in order to figure out what changes to make. They figure out what changes to make. They figure out what changes to make. They can grow their own context over time and

  11. can grow their own context over time and can grow their own context over time and prompt themselves with sub agents. They prompt themselves with sub agents. They prompt themselves with sub agents. They can do all of these different things. It can do all of these different things. It can do all of these different things. It makes no sense at all to give a model makes no sense at all to give a model makes no sense at all to give a model that takes a bit of text input and that takes a bit of text input and that takes a bit of text input and immediately responds with JSON three immediately responds with JSON three immediately responds with JSON three different implementations of something different implementations of something different implementations of something by three different LMS. It just doesn't by three different LMS. It just doesn't by three different LMS. It just doesn't know enough to make a [snorts] good know enough to make a [snorts] good know enough to make a [snorts] good decision there. I want to beat into your decision there. I want to beat into your decision there. I want to beat into your guys' heads how this works. And I'm guys' heads how this works. And I'm guys' heads how this works. And I'm sorry for those who get it because it's sorry for those who get it because it's sorry for those who get it because it's going to be tedious, but I'm going to be tedious, but I'm going to be tedious, but I'm increasingly tired of comment sections increasingly tired of comment sections increasingly tired of comment sections that fundamentally don't understand that fundamentally don't understand that fundamentally don't understand what's being talked about. This is a what's being talked about. This is a what's being talked about. This is a model that works like a switch model that works like a switch model that works like a switch statement. It is roughly as intelligent statement. It is roughly as intelligent statement. It is roughly as intelligent as a switch statement. It's for as a switch statement. It's for as a switch statement. It's for classifying things. It doesn't know how classifying things. It doesn't know how classifying things. It doesn't know how to look through a codebase to make a to look through a codebase to make a to look through a codebase to make a good decision. It doesn't have the good decision. It doesn't have the good decision. It doesn't have the intelligence to distinguish between the intelligence to distinguish between the intelligence to distinguish between the outputs of language models. And most outputs of language models. And most outputs of language models. And most importantly, the context window is tiny. importantly, the context window is tiny. importantly, the context window is tiny. It's only 32k tokens. That doesn't mean It's only 32k tokens. That doesn't mean It's only 32k tokens. That doesn't mean it's not cool, and it is cool as hell. it's not cool, and it is cool as hell. it's not cool, and it is cool as hell. The speed and the cost is insane. They The speed and the cost is insane. They The speed and the cost is insane. They do a comparison here against an LLM for do a comparison here against an LLM for do a comparison here against an LLM for responding to a query. responding to a query. responding to a query. Type safe race ask LM versus type safe. Type safe race ask LM versus type safe. Type safe race ask LM versus type safe. He got back the response immediately. It He got back the response immediately. It He got back the response immediately. It was given 27 questions about a bunch of was given 27 questions about a bunch of was given 27 questions about a bunch of data and it responded to all of them data and it responded to all of them data and it responded to all of them with the right format with numbers with the right format with numbers with the right format with numbers ranking them and the type as well as ranking them and the type as well as ranking them and the type as well as confidence on its decisions for these confidence on its decisions for these confidence on its decisions for these things. It took 0.114 seconds and it things. It took 0.114 seconds and it things. It took 0.114 seconds and it cost a amount that rounds to zero.

  12. cost a amount that rounds to zero. cost a amount that rounds to zero. Meanwhile, Terra once it finally Meanwhile, Terra once it finally Meanwhile, Terra once it finally finished took 9 seconds and cost 1.3. finished took 9 seconds and cost 1.3. finished took 9 seconds and cost 1.3. That is a comical gap here. 170x cheaper That is a comical gap here. 170x cheaper That is a comical gap here. 170x cheaper and 70s something times faster. Respect and 70s something times faster. Respect and 70s something times faster. Respect to them for keeping the cost per to them for keeping the cost per to them for keeping the cost per workflow in their chart in log. They workflow in their chart in log. They workflow in their chart in log. They could have made this linear and it would could have made this linear and it would could have made this linear and it would have looked hilarious. The only models have looked hilarious. The only models have looked hilarious. The only models they measured that had better they measured that had better they measured that had better classification in their demo workflows classification in their demo workflows classification in their demo workflows than Jev were soul and opus 5. They than Jev were soul and opus 5. They than Jev were soul and opus 5. They didn't test Fable or Astra, but like didn't test Fable or Astra, but like didn't test Fable or Astra, but like those are way too expensive. You those are way too expensive. You those are way too expensive. You shouldn't even look at them for this. shouldn't even look at them for this. shouldn't even look at them for this. But it is outranking DSV4 Flash in But it is outranking DSV4 Flash in But it is outranking DSV4 Flash in various tests for ranking things. and it various tests for ranking things. and it various tests for ranking things. and it owns the PTO Frontier with almost two owns the PTO Frontier with almost two owns the PTO Frontier with almost two orders of magnitude again because it's orders of magnitude again because it's orders of magnitude again because it's very specifically focused on this one very specifically focused on this one very specifically focused on this one thing. They didn't publish the bench thing. They didn't publish the bench thing. They didn't publish the bench they used here, but they gave some they used here, but they gave some they used here, but they gave some examples. There's an alert. The next examples. There's an alert. The next examples. There's an alert. The next step is a question. Is this unauthorized step is a question. Is this unauthorized step is a question. Is this unauthorized activity with options for what it can activity with options for what it can activity with options for what it can choose? Once that happens, we determine choose? Once that happens, we determine choose? Once that happens, we determine different flows it goes through. And if different flows it goes through. And if different flows it goes through. And if we have decided that this is an we have decided that this is an we have decided that this is an incident, then we ask what state is the incident, then we ask what state is the incident, then we ask what state is the incident in and it has 11 readings for incident in and it has 11 readings for incident in and it has 11 readings for it. It processes all the data. We then it. It processes all the data. We then it. It processes all the data. We then ask what action it thinks we should take ask what action it thinks we should take ask what action it thinks we should take and then we rank the results after. They and then we rank the results after. They and then we rank the results after. They are surprisingly transparent with a lot are surprisingly transparent with a lot are surprisingly transparent with a lot of things. Like they love to say all the of things. Like they love to say all the of things. Like they love to say all the things the model isn't good at. This is things the model isn't good at. This is things the model isn't good at. This is where the numbers on their homepage come where the numbers on their homepage come where the numbers on their homepage come from. The 193.6x faster and 444.6x from. The 193.6x faster and 444.6x from. The 193.6x faster and 444.6x cheaper. But they say that they expect cheaper. But they say that they expect cheaper. But they say that they expect this is the higher end of real world this is the higher end of real world this is the higher end of real world gains. The context of these workflows gains. The context of these workflows gains. The context of these workflows were not deliberately chosen nor were not deliberately chosen nor were not deliberately chosen nor constructed to make our model look good.

  13. constructed to make our model look good. constructed to make our model look good. Cool. They're using the average of GBD6 Cool. They're using the average of GBD6 Cool. They're using the average of GBD6 Astra and Fable 5.1 as the reference Astra and Fable 5.1 as the reference Astra and Fable 5.1 as the reference answer. That's why they don't appear in answer. That's why they don't appear in answer. That's why they don't appear in the list and also why the data might not the list and also why the data might not the list and also why the data might not be perfect since those models might be be perfect since those models might be be perfect since those models might be wrong as well. The LMS use their system wrong as well. The LMS use their system wrong as well. The LMS use their system 1 LM wrapper which constrains LLMs to 1 LM wrapper which constrains LLMs to 1 LM wrapper which constrains LLMs to output structured decisions compatible output structured decisions compatible output structured decisions compatible with their API. We found it to be the with their API. We found it to be the with their API. We found it to be the most accurate way to get decisions from most accurate way to get decisions from most accurate way to get decisions from LLMs, but this tends to be slower and LLMs, but this tends to be slower and LLMs, but this tends to be slower and more expensive than giving decisions more expensive than giving decisions more expensive than giving decisions without probabilities. I find that they without probabilities. I find that they without probabilities. I find that they love to tie the idea of hallucination love to tie the idea of hallucination love to tie the idea of hallucination and type safety. Type safety being you and type safety. Type safety being you and type safety. Type safety being you have a contract for what it should have a contract for what it should have a contract for what it should output and it follows the contract. output and it follows the contract. output and it follows the contract. That's why they're type- safe AI because That's why they're type- safe AI because That's why they're type- safe AI because the shape of the data will always be the shape of the data will always be the shape of the data will always be honored. Hallucination goes a lot honored. Hallucination goes a lot honored. Hallucination goes a lot further than hallucinating fields in a further than hallucinating fields in a further than hallucinating fields in a return type. But in this particular return type. But in this particular return type. But in this particular case, it can be pretty brutal if your LM case, it can be pretty brutal if your LM case, it can be pretty brutal if your LM hallucinates that it can change the hallucinates that it can change the hallucinates that it can change the fields when it can't. That does suck for fields when it can't. That does suck for fields when it can't. That does suck for type safety. The way they frame it here type safety. The way they frame it here type safety. The way they frame it here is that having a hallucinated tool call is that having a hallucinated tool call is that having a hallucinated tool call can be inconvenient for an agent, but can be inconvenient for an agent, but can be inconvenient for an agent, but it's an absolute dealbreaker if it's it's an absolute dealbreaker if it's it's an absolute dealbreaker if it's part of a system with latency guarantees part of a system with latency guarantees part of a system with latency guarantees or if it's buried several layers deep in or if it's buried several layers deep in or if it's buried several layers deep in a dependency chain. Yep, that part I a dependency chain. Yep, that part I a dependency chain. Yep, that part I absolutely agree with. No matter how absolutely agree with. No matter how absolutely agree with. No matter how smart an existing model is, it can still smart an existing model is, it can still smart an existing model is, it can still hallucinate and have type errors. They hallucinate and have type errors. They hallucinate and have type errors. They show that here where crazy enough, Astra show that here where crazy enough, Astra show that here where crazy enough, Astra actually has more errors with structured actually has more errors with structured actually has more errors with structured tool outputs than sole terra and Luna.

  14. tool outputs than sole terra and Luna. tool outputs than sole terra and Luna. It actually has more than terror and It actually has more than terror and It actually has more than terror and Luna combined, which is kind of crazy. Luna combined, which is kind of crazy. Luna combined, which is kind of crazy. But then you start to look at models But then you start to look at models But then you start to look at models from a certain anthropic where haiku had from a certain anthropic where haiku had from a certain anthropic where haiku had a 45.5% a 45.5% a 45.5% error rate. That model is so bad and it error rate. That model is so bad and it error rate. That model is so bad and it needs to stop being used for anything. needs to stop being used for anything. needs to stop being used for anything. It's about a year old now might be over It's about a year old now might be over It's about a year old now might be over that. Just anthropic is treating haiku that. Just anthropic is treating haiku that. Just anthropic is treating haiku as dead. We should do them the favor of as dead. We should do them the favor of as dead. We should do them the favor of doing the same. But of course the number doing the same. But of course the number doing the same. But of course the number that matters here is jev with a 0%. This that matters here is jev with a 0%. This that matters here is jev with a 0%. This is for structured output errors. When is for structured output errors. When is for structured output errors. When you give it a JSON format and tell it to you give it a JSON format and tell it to you give it a JSON format and tell it to honor it, Haiku cannot do it. If we honor it, Haiku cannot do it. If we honor it, Haiku cannot do it. If we switch over to tool call error rates, switch over to tool call error rates, switch over to tool call error rates, funny enough, anthrowing models start funny enough, anthrowing models start funny enough, anthrowing models start doing a lot better and OpenAI models doing a lot better and OpenAI models doing a lot better and OpenAI models start to have more problems. But of start to have more problems. But of start to have more problems. But of course, Jev still is at zero. It honors course, Jev still is at zero. It honors course, Jev still is at zero. It honors the format it's given. They made a demo the format it's given. They made a demo the format it's given. They made a demo of it playing Doom and the engineer who of it playing Doom and the engineer who of it playing Doom and the engineer who made it was concerned about cost, but made it was concerned about cost, but made it was concerned about cost, but they figured out even though it's they figured out even though it's they figured out even though it's running 10 times a second, the cost running 10 times a second, the cost running 10 times a second, the cost would still be under $7 an hour because would still be under $7 an hour because would still be under $7 an hour because of how efficient the model is and how of how efficient the model is and how of how efficient the model is and how cheap it is to run. Remember, it doesn't cheap it is to run. Remember, it doesn't cheap it is to run. Remember, it doesn't have vision. So this is all from game have vision. So this is all from game have vision. So this is all from game state data that it's being given. The state data that it's being given. The state data that it's being given. The application state is handed to the model application state is handed to the model application state is handed to the model and it's given the JSON format on what and it's given the JSON format on what and it's given the JSON format on what it should decide to do next and it it should decide to do next and it it should decide to do next and it decides. You will see some quirks with decides. You will see some quirks with decides. You will see some quirks with that though. Watch how often it like that though. Watch how often it like that though. Watch how often it like just like rotates left and right wildly.

  15. just like rotates left and right wildly. just like rotates left and right wildly. That's because every single frame it's That's because every single frame it's That's because every single frame it's told make a decision. So it doesn't told make a decision. So it doesn't told make a decision. So it doesn't necessarily know what decision it made necessarily know what decision it made necessarily know what decision it made before. It doesn't know it turned left before. It doesn't know it turned left before. It doesn't know it turned left so keep turning left. So it's like okay so keep turning left. So it's like okay so keep turning left. So it's like okay looking here I'll turn left. Okay state looking here I'll turn left. Okay state looking here I'll turn left. Okay state turn right. on every single frame turn right. on every single frame turn right. on every single frame effectively its brain is being wiped and effectively its brain is being wiped and effectively its brain is being wiped and it's making a new decision. You get the it's making a new decision. You get the it's making a new decision. You get the idea though and I saw this with the idea though and I saw this with the idea though and I saw this with the chess as well. It's not meant for chess as well. It's not meant for chess as well. It's not meant for longunning jobs. It's meant for making a longunning jobs. It's meant for making a longunning jobs. It's meant for making a quick decision on the fly or being part quick decision on the fly or being part quick decision on the fly or being part of some other longunning jobs work. They of some other longunning jobs work. They of some other longunning jobs work. They call out that this demo is again on call out that this demo is again on call out that this demo is again on structured state as a data structure structured state as a data structure structured state as a data structure with text not on images yet. This is with text not on images yet. This is with text not on images yet. This is particularly exciting. This model having particularly exciting. This model having particularly exciting. This model having image support will be super super image support will be super super image support will be super super useful. The model is also insane at wiki useful. The model is also insane at wiki useful. The model is also insane at wiki racing because it makes decisions so racing because it makes decisions so racing because it makes decisions so quick. So when it has the contents of an quick. So when it has the contents of an quick. So when it has the contents of an HTML page and it knows where the links HTML page and it knows where the links HTML page and it knows where the links are on it, it can choose which one to are on it, it can choose which one to are on it, it can choose which one to click and continue going much faster. Yeah, it is very very fast. Terra even Yeah, it is very very fast. Terra even had a hallucination during its run.

  16. had a hallucination during its run. had a hallucination during its run. That's funny. Let's take a look at what That's funny. Let's take a look at what That's funny. Let's take a look at what people are using the model for. They people are using the model for. They people are using the model for. They claim you can't use it for generating claim you can't use it for generating claim you can't use it for generating text, but if you give it the ability to text, but if you give it the ability to text, but if you give it the ability to respond by choosing which character it respond by choosing which character it respond by choosing which character it thinks makes the most sense, you can thinks makes the most sense, you can thinks makes the most sense, you can kind of get it to do things. Here's one kind of get it to do things. Here's one kind of get it to do things. Here's one that I think actually makes way more that I think actually makes way more that I think actually makes way more sense from Chris over at Verscell. It's sense from Chris over at Verscell. It's sense from Chris over at Verscell. It's the idea of using JSON render with Jev. the idea of using JSON render with Jev. the idea of using JSON render with Jev. The point of JSON render is to make it The point of JSON render is to make it The point of JSON render is to make it easy to create a component and use it easy to create a component and use it easy to create a component and use it for arbitrary JSON. So you can link the for arbitrary JSON. So you can link the for arbitrary JSON. So you can link the outputs of an LLM into your UI in a more outputs of an LLM into your UI in a more outputs of an LLM into your UI in a more visible and useful way, but you have to visible and useful way, but you have to visible and useful way, but you have to wait for the LM to generate the JSON wait for the LM to generate the JSON wait for the LM to generate the JSON before you can update the UI. What if before you can update the UI. What if before you can update the UI. What if the model could do that comically the model could do that comically the model could do that comically faster? That's a pretty big difference. That's a pretty big difference. Yeah, it's insane. It responds with the Yeah, it's insane. It responds with the Yeah, it's insane. It responds with the whole thing at once because it's not whole thing at once because it's not whole thing at once because it's not streaming in text. Everybody's been streaming in text. Everybody's been streaming in text. Everybody's been telling me Ryan's been cooking with it. telling me Ryan's been cooking with it. telling me Ryan's been cooking with it. Let's see what he's done. 1,500 of my Let's see what he's done. 1,500 of my Let's see what he's done. 1,500 of my own emails that I've exported from uh my own emails that I've exported from uh my own emails that I've exported from uh my email. And uh we're going to run this email. And uh we're going to run this email. And uh we're going to run this classification model because I got classification model because I got classification model because I got access to it and see how well it runs access to it and see how well it runs access to it and see how well it runs the classification on it. We're going to the classification on it. We're going to the classification on it. We're going to do a batch of 100 emails just to start do a batch of 100 emails just to start do a batch of 100 emails just to start out with eight workers and we'll see how out with eight workers and we'll see how out with eight workers and we'll see how long it takes to do it. So, let's start.

  17. long it takes to do it. So, let's start. long it takes to do it. So, let's start. Boom. 100 emails done. It doesn't tell Boom. 100 emails done. It doesn't tell Boom. 100 emails done. It doesn't tell me how long it took. average was 200 me how long it took. average was 200 me how long it took. average was 200 milliseconds. milliseconds. milliseconds. 200 milliseconds per email. P95 was 240 200 milliseconds per email. P95 was 240 200 milliseconds per email. P95 was 240 milliseconds and it did 38 per second, milliseconds and it did 38 per second, milliseconds and it did 38 per second, which is amazing. So, this is a much which is amazing. So, this is a much which is amazing. So, this is a much better example. I already use LLMs to go better example. I already use LLMs to go better example. I already use LLMs to go through my email. It's expensive, but through my email. It's expensive, but through my email. It's expensive, but it's fine. This is a first pass to like it's fine. This is a first pass to like it's fine. This is a first pass to like get the quick, loweffort spam things out get the quick, loweffort spam things out get the quick, loweffort spam things out of your way. Super useful. It's the of your way. Super useful. It's the of your way. Super useful. It's the things that are too cheap to justify things that are too cheap to justify things that are too cheap to justify running an LM for or the things that are running an LM for or the things that are running an LM for or the things that are run too often to wait that much time run too often to wait that much time run too often to wait that much time for. Once this has image support, an for. Once this has image support, an for. Once this has image support, an example of something we can do with this example of something we can do with this example of something we can do with this is take a bunch of frames for my video is take a bunch of frames for my video is take a bunch of frames for my video and ask it, does this frame have and ask it, does this frame have and ask it, does this frame have sensitive data or PII on it, like does sensitive data or PII on it, like does sensitive data or PII on it, like does it have an email address on it? And to it have an email address on it? And to it have an email address on it? And to tell me where it is some amount so I can tell me where it is some amount so I can tell me where it is some amount so I can go find these and clean up our videos go find these and clean up our videos go find these and clean up our videos before we release them. That type of before we release them. That type of before we release them. That type of thing is so nice. I think computer use thing is so nice. I think computer use thing is so nice. I think computer use will also be way more compelling with it will also be way more compelling with it will also be way more compelling with it once it has the ability to see what's once it has the ability to see what's once it has the ability to see what's going on, but it's already really good going on, but it's already really good going on, but it's already really good at navigating websites because it can at navigating websites because it can at navigating websites because it can just take the HTML and then decide what just take the HTML and then decide what just take the HTML and then decide what to do based on the current page content.

  18. to do based on the current page content. to do based on the current page content. There's an example of it booking flights There's an example of it booking flights There's an example of it booking flights in under 7.1 seconds, right? Yeah, in under 7.1 seconds, right? Yeah, in under 7.1 seconds, right? Yeah, that's crazy. LM's doing this take so that's crazy. LM's doing this take so that's crazy. LM's doing this take so much longer. I am scared to even make much longer. I am scared to even make much longer. I am scared to even make this video if I'm being real because if this video if I'm being real because if this video if I'm being real because if I increase excitement too much around I increase excitement too much around I increase excitement too much around this model, we're going to end up with this model, we're going to end up with this model, we're going to end up with some really, really dumb things. For some really, really dumb things. For some really, really dumb things. For example, this tweet from Brain Trust. example, this tweet from Brain Trust. example, this tweet from Brain Trust. Their goal is to increase observability Their goal is to increase observability Their goal is to increase observability for agents. This is a really bad thing for agents. This is a really bad thing for agents. This is a really bad thing for them to tweet. It makes me not trust for them to tweet. It makes me not trust for them to tweet. It makes me not trust them because they suggest that you them because they suggest that you them because they suggest that you should use Jev as a judge scorer in a should use Jev as a judge scorer in a should use Jev as a judge scorer in a tool like Brain Trust to decide between tool like Brain Trust to decide between tool like Brain Trust to decide between outputs from LM. They're legitimately outputs from LM. They're legitimately outputs from LM. They're legitimately suggesting that you replace an LLM that suggesting that you replace an LLM that suggesting that you replace an LLM that you use for judging for scoring agent you use for judging for scoring agent you use for judging for scoring agent responses with Jev. And the reason is so responses with Jev. And the reason is so responses with Jev. And the reason is so that you don't need to spend time and that you don't need to spend time and that you don't need to spend time and resources prompting a general purpose resources prompting a general purpose resources prompting a general purpose model into an LLM judge. What those model into an LLM judge. What those model into an LLM judge. What those prompts take 30 seconds to prompts take 30 seconds to prompts take 30 seconds to write. Not even. And like it's if you write. Not even. And like it's if you write. Not even. And like it's if you just give it the examples from something just give it the examples from something just give it the examples from something like Jev, it's going to follow them. like Jev, it's going to follow them. like Jev, it's going to follow them. This is hilarious. On that note, there This is hilarious. On that note, there This is hilarious. On that note, there was another demo that I actually think was another demo that I actually think was another demo that I actually think is really cool conceptually. like it is really cool conceptually. like it is really cool conceptually. like it looks crazy and it helps show what looks crazy and it helps show what looks crazy and it helps show what capabilities this model has, but if you capabilities this model has, but if you capabilities this model has, but if you actually think it's a good idea to use a actually think it's a good idea to use a actually think it's a good idea to use a non-reasoning classifier model for non-reasoning classifier model for non-reasoning classifier model for compaction for your context, then I compaction for your context, then I compaction for your context, then I would plead that you never ever ever would plead that you never ever ever would plead that you never ever ever stray from the defaults in those tools stray from the defaults in those tools stray from the defaults in those tools because you just fundamentally don't because you just fundamentally don't because you just fundamentally don't understand yet. And don't worry, you're understand yet. And don't worry, you're understand yet. And don't worry, you're not the only one. I don't think anyone not the only one. I don't think anyone not the only one. I don't think anyone has to. Context compaction is a complex has to. Context compaction is a complex has to. Context compaction is a complex topic and it's gotten more complex over

  19. topic and it's gotten more complex over topic and it's gotten more complex over the years. I will do my best to tell the years. I will do my best to tell the years. I will do my best to tell TLDDR why this is a bad idea, but you TLDDR why this is a bad idea, but you TLDDR why this is a bad idea, but you should just read my longer post if should just read my longer post if should just read my longer post if you're curious. First thing, compaction you're curious. First thing, compaction you're curious. First thing, compaction isn't a filter. We're not just going isn't a filter. We're not just going isn't a filter. We're not just going through your history and selectively through your history and selectively through your history and selectively deleting lines from it. We are deleting lines from it. We are deleting lines from it. We are synthesizing a summary based on synthesizing a summary based on synthesizing a summary based on everything that's happened so far. Jev everything that's happened so far. Jev everything that's happened so far. Jev also doesn't have enough context to even also doesn't have enough context to even also doesn't have enough context to even know what it's deciding on. Then we just know what it's deciding on. Then we just know what it's deciding on. Then we just saw it doesn't have the tool called saw it doesn't have the tool called saw it doesn't have the tool called outputs and results. It only has the outputs and results. It only has the outputs and results. It only has the inputs and a little bit of the context inputs and a little bit of the context inputs and a little bit of the context from the thread and not that much of it from the thread and not that much of it from the thread and not that much of it cuz it's 32k tokens of context. So it cuz it's 32k tokens of context. So it cuz it's 32k tokens of context. So it doesn't have enough data to make a doesn't have enough data to make a doesn't have enough data to make a decision. Even separate from the fact decision. Even separate from the fact decision. Even separate from the fact that it doesn't have access to the that it doesn't have access to the that it doesn't have access to the reasoning data at all because the reasoning data at all because the reasoning data at all because the reasoning data is never shared by the reasoning data is never shared by the reasoning data is never shared by the labs anymore. When you call the claude labs anymore. When you call the claude labs anymore. When you call the claude or codeex APIs, you don't get back or codeex APIs, you don't get back or codeex APIs, you don't get back reasoning. You might get back an reasoning. You might get back an reasoning. You might get back an encrypted payload that they can then map encrypted payload that they can then map encrypted payload that they can then map to the reasoning on their end. Or you to the reasoning on their end. Or you to the reasoning on their end. Or you might get a summary if you're lucky, but might get a summary if you're lucky, but might get a summary if you're lucky, but you don't know what the model was you don't know what the model was you don't know what the model was thinking when it made a decision. So any thinking when it made a decision. So any thinking when it made a decision. So any attempt to compact that is not going to attempt to compact that is not going to attempt to compact that is not going to include those decisions. This gets even include those decisions. This gets even include those decisions. This gets even worse when you remember that Enthropic worse when you remember that Enthropic worse when you remember that Enthropic is making changes to how history is making changes to how history is making changes to how history preservation works such that if you edit preservation works such that if you edit preservation works such that if you edit your history, you lose all of the your history, you lose all of the your history, you lose all of the reasoning traces for that context.

  20. reasoning traces for that context. reasoning traces for that context. There's also the fact that models are There's also the fact that models are There's also the fact that models are tuned on the way that they compact. They tuned on the way that they compact. They tuned on the way that they compact. They do this in training. Now, models learn do this in training. Now, models learn do this in training. Now, models learn how to handle compaction well and they how to handle compaction well and they how to handle compaction well and they make adjustments to the weights and how make adjustments to the weights and how make adjustments to the weights and how this works through the process of this works through the process of this works through the process of training through RL. The compaction that training through RL. The compaction that training through RL. The compaction that the models do is already pretty damn the models do is already pretty damn the models do is already pretty damn good and changing what history they have good and changing what history they have good and changing what history they have before they do it makes even less sense. before they do it makes even less sense. before they do it makes even less sense. Another important thing to recognize is Another important thing to recognize is Another important thing to recognize is that cache writes are often, if not that cache writes are often, if not that cache writes are often, if not always, for agentic use cases, much more always, for agentic use cases, much more always, for agentic use cases, much more expensive than cash reads end up being. expensive than cash reads end up being. expensive than cash reads end up being. And when you remember how cash And when you remember how cash And when you remember how cash invalidation works, you realize that invalidation works, you realize that invalidation works, you realize that this will probably break the cache quite this will probably break the cache quite this will probably break the cache quite a bit if you run it actively enough. a bit if you run it actively enough. a bit if you run it actively enough. Because if you have a history like 1 2 3 Because if you have a history like 1 2 3 Because if you have a history like 1 2 3 4 5 6 and then you delete number two, 4 5 6 and then you delete number two, 4 5 6 and then you delete number two, everything three onwards has to be everything three onwards has to be everything three onwards has to be rewritten because the history has to be rewritten because the history has to be rewritten because the history has to be prefixed. Any changes early mean prefixed. Any changes early mean prefixed. Any changes early mean everything past that point's everything past that point's everything past that point's invalidated. The model's given weird invalidated. The model's given weird invalidated. The model's given weird instructions as implementation, too. instructions as implementation, too. instructions as implementation, too. Things like whatever's not kept is Things like whatever's not kept is Things like whatever's not kept is permanently deleted, but the assistant permanently deleted, but the assistant permanently deleted, but the assistant can rerun a tool if needed. Some tools can rerun a tool if needed. Some tools can rerun a tool if needed. Some tools are destructive. It's not that simple. are destructive. It's not that simple. are destructive. It's not that simple. And you'll also get to a point if it's And you'll also get to a point if it's And you'll also get to a point if it's classifying in such a shallow way where classifying in such a shallow way where classifying in such a shallow way where it gets stuck in a loop where it's it gets stuck in a loop where it's it gets stuck in a loop where it's already removed everything it thinks already removed everything it thinks already removed everything it thinks doesn't matter and it only has left what doesn't matter and it only has left what doesn't matter and it only has left what does. No. And people have actually tried does. No. And people have actually tried does. No. And people have actually tried this and benched it. It doesn't perform this and benched it. It doesn't perform this and benched it. It doesn't perform well at all. There is one benefit to well at all. There is one benefit to well at all. There is one benefit to this style of compaction. It's so fast this style of compaction. It's so fast this style of compaction. It's so fast that it fits in the attention span of that it fits in the attention span of that it fits in the attention span of the average Twitter user. So, the video the average Twitter user. So, the video the average Twitter user. So, the video is guaranteed to go viral. But you're is guaranteed to go viral. But you're is guaranteed to go viral. But you're not like the average Twitter user.

  21. not like the average Twitter user. not like the average Twitter user. You've been watching this video for much You've been watching this video for much You've been watching this video for much longer than the attention span of the longer than the attention span of the longer than the attention span of the average Twitter user. And for that, I average Twitter user. And for that, I average Twitter user. And for that, I appreciate you. If you haven't hit the appreciate you. If you haven't hit the appreciate you. If you haven't hit the sub button, I would appreciate that as sub button, I would appreciate that as sub button, I would appreciate that as well, cuz a lot of y'all haven't. And it well, cuz a lot of y'all haven't. And it well, cuz a lot of y'all haven't. And it seems like you want this type of long- seems like you want this type of long- seems like you want this type of long- form content. You should consider form content. You should consider form content. You should consider subscribing to signify that. And as a subscribing to signify that. And as a subscribing to signify that. And as a Twitter user, trust me, you don't want Twitter user, trust me, you don't want Twitter user, trust me, you don't want to be like us. Avoid it to the best of to be like us. Avoid it to the best of to be like us. Avoid it to the best of your ability. I do think there are ways your ability. I do think there are ways your ability. I do think there are ways that a tool like Jev can be useful to that a tool like Jev can be useful to that a tool like Jev can be useful to things like agentic dev work. Not having things like agentic dev work. Not having things like agentic dev work. Not having it write code or compact by context or it write code or compact by context or it write code or compact by context or change anything that the harness is change anything that the harness is change anything that the harness is doing. The harnesses are pretty dang doing. The harnesses are pretty dang doing. The harnesses are pretty dang good now. You should lean into the fact good now. You should lean into the fact good now. You should lean into the fact that billions upon billions of dollars that billions upon billions of dollars that billions upon billions of dollars are being spent on that and not reinvent are being spent on that and not reinvent are being spent on that and not reinvent the wheel constantly. Don't get me the wheel constantly. Don't get me the wheel constantly. Don't get me started on the people who think they can started on the people who think they can started on the people who think they can compact their history by taking their compact their history by taking their compact their history by taking their entire context and then shoving it into entire context and then shoving it into entire context and then shoving it into a small image. Insanity. Anyways, a a small image. Insanity. Anyways, a a small image. Insanity. Anyways, a thing you can use this for is going thing you can use this for is going thing you can use this for is going through large amounts of data. For through large amounts of data. For through large amounts of data. For example, all of your history using example, all of your history using example, all of your history using models. Here are my 1,118T3 models. Here are my 1,118T3 models. Here are my 1,118T3 code chats that I did on this particular code chats that I did on this particular code chats that I did on this particular machine categorized and classified using machine categorized and classified using machine categorized and classified using Jev. It classified 32,311 Jev. It classified 32,311 Jev. It classified 32,311 messages across 1118 threads. And doing messages across 1118 threads. And doing messages across 1118 threads. And doing all of that cost 37.

  22. all of that cost 37. all of that cost 37. And it found that nearly half of what I And it found that nearly half of what I And it found that nearly half of what I did was bug fixing. and PR work, did was bug fixing. and PR work, did was bug fixing. and PR work, specifically reviewing and managing PR. specifically reviewing and managing PR. specifically reviewing and managing PR. So was 20% of my threads. Remember that So was 20% of my threads. Remember that So was 20% of my threads. Remember that the results here are classifications. So the results here are classifications. So the results here are classifications. So they're usually scored. It's not they're usually scored. It's not they're usually scored. It's not responding with a list of strings on responding with a list of strings on responding with a list of strings on what to tag. It's taking all of the what to tag. It's taking all of the what to tag. It's taking all of the values inside of the body there and values inside of the body there and values inside of the body there and giving you a threshold on a scale of 0 giving you a threshold on a scale of 0 giving you a threshold on a scale of 0 to one for all of them. And you can to one for all of them. And you can to one for all of them. And you can change what you want to count and not change what you want to count and not change what you want to count and not count. For example, here with the simple count. For example, here with the simple count. For example, here with the simple web app I made, I am changing what bar web app I made, I am changing what bar web app I made, I am changing what bar we have set for each of the we have set for each of the we have set for each of the classifications. So if we have it at an classifications. So if we have it at an classifications. So if we have it at an 80% confidence interval, then 22% of my 80% confidence interval, then 22% of my 80% confidence interval, then 22% of my threads are expanding scope. But if I threads are expanding scope. But if I threads are expanding scope. But if I bump that up to 90%, it's down to 6.8 bump that up to 90%, it's down to 6.8 bump that up to 90%, it's down to 6.8 because it was not sure in the rest of because it was not sure in the rest of because it was not sure in the rest of those cases or as sure so to speak. those cases or as sure so to speak. those cases or as sure so to speak. Funny enough, when I had Astra build Funny enough, when I had Astra build Funny enough, when I had Astra build this, it did note some of the things this, it did note some of the things this, it did note some of the things that it wasn't great at, particularly that it wasn't great at, particularly that it wasn't great at, particularly trying to identify which threads could trying to identify which threads could trying to identify which threads could be used for content for me. So, one of be used for content for me. So, one of be used for content for me. So, one of the ideas it had was what if we can the ideas it had was what if we can the ideas it had was what if we can figure out which threads Theo should figure out which threads Theo should figure out which threads Theo should save to use in a video. And it tried save to use in a video. And it tried save to use in a video. And it tried multiple times to revise the prompt to multiple times to revise the prompt to multiple times to revise the prompt to get better threads for that. And it get better threads for that. And it get better threads for that. And it still concluded about half of my threads still concluded about half of my threads still concluded about half of my threads were worth using for content. So, it's were worth using for content. So, it's were worth using for content. So, it's not good at this. It's not smart enough not good at this. It's not smart enough not good at this. It's not smart enough to do those types of complex reasoning to do those types of complex reasoning to do those types of complex reasoning things whereas to make a a decision that things whereas to make a a decision that things whereas to make a a decision that requires thought. It's like one last requires thought. It's like one last requires thought. It's like one last simple way to think of this. How many simple way to think of this. How many simple way to think of this. How many seconds would it take for you to answer seconds would it take for you to answer seconds would it take for you to answer the question once you've perceived all the question once you've perceived all the question once you've perceived all of the information? If I show you a

  23. of the information? If I show you a of the information? If I show you a picture of a person wearing a shirt, how picture of a person wearing a shirt, how picture of a person wearing a shirt, how many seconds does it take for you to say many seconds does it take for you to say many seconds does it take for you to say what color it is? If the answer is under what color it is? If the answer is under what color it is? If the answer is under 10 seconds, this model's probably good 10 seconds, this model's probably good 10 seconds, this model's probably good for it. If it's over 10 seconds, this for it. If it's over 10 seconds, this for it. If it's over 10 seconds, this model's likely less good for it. This is model's likely less good for it. This is model's likely less good for it. This is the point I'll end on, and it's similar the point I'll end on, and it's similar the point I'll end on, and it's similar to the one I started on. really want to to the one I started on. really want to to the one I started on. really want to emphasize the system one point because emphasize the system one point because emphasize the system one point because I'm tired of people not getting it. If I'm tired of people not getting it. If I'm tired of people not getting it. If the task requires thinking, this model the task requires thinking, this model the task requires thinking, this model is not right. If the task requires is not right. If the task requires is not right. If the task requires classifying, organizing, ranking, real classifying, organizing, ranking, real classifying, organizing, ranking, real quick decision-m, this model is quick decision-m, this model is quick decision-m, this model is incredible. It's a system one model. incredible. It's a system one model. incredible. It's a system one model. System two is when you think about a System two is when you think about a System two is when you think about a thing. System one is when you react thing. System one is when you react thing. System one is when you react because your brain thought for you. It's because your brain thought for you. It's because your brain thought for you. It's almost like the reflex. It's like when almost like the reflex. It's like when almost like the reflex. It's like when you tap your knee and your leg kicks you tap your knee and your leg kicks you tap your knee and your leg kicks out, that kind of thing. This model is out, that kind of thing. This model is out, that kind of thing. This model is designed to work like that first part to designed to work like that first part to designed to work like that first part to be really fast and quick. If you're be really fast and quick. If you're be really fast and quick. If you're thinking of this model in terms of how thinking of this model in terms of how thinking of this model in terms of how it replaces the other ones you use, it replaces the other ones you use, it replaces the other ones you use, you're probably not thinking about it you're probably not thinking about it you're probably not thinking about it correctly, unless you're doing a lot of correctly, unless you're doing a lot of correctly, unless you're doing a lot of structured output work. The value of structured output work. The value of structured output work. The value of this model is that it made so many this model is that it made so many this model is that it made so many things that weren't really realistic things that weren't really realistic things that weren't really realistic before way cheaper and faster in a way before way cheaper and faster in a way before way cheaper and faster in a way that is actually kind of cool as a that is actually kind of cool as a that is actually kind of cool as a programmer. But you shouldn't be viewing programmer. But you shouldn't be viewing programmer. But you shouldn't be viewing this as a tool you use in Codeex or this as a tool you use in Codeex or this as a tool you use in Codeex or Claude or even in your terminal. You Claude or even in your terminal. You Claude or even in your terminal. You should see this like a new library you should see this like a new library you should see this like a new library you install or a function that you call. It install or a function that you call. It install or a function that you call. It is meant to be integrated in the tools is meant to be integrated in the tools is meant to be integrated in the tools that we build, not used as an inference that we build, not used as an inference that we build, not used as an inference farm to generate code or text or all farm to generate code or text or all farm to generate code or text or all these other things. Somebody in China these other things. Somebody in China these other things. Somebody in China just said the only way this would just said the only way this would just said the only way this would replace an LM is if you're using them replace an LM is if you're using them replace an LM is if you're using them badly. I mostly agree. There are a lot badly. I mostly agree. There are a lot badly. I mostly agree. There are a lot of use case for structured output type of use case for structured output type of use case for structured output type stuff and this is where it is strongest.

  24. stuff and this is where it is strongest. stuff and this is where it is strongest. I think this model is really cool and I think this model is really cool and I think this model is really cool and I've been enjoying it a lot. I have a I've been enjoying it a lot. I have a I've been enjoying it a lot. I have a feeling you will too as long as you go feeling you will too as long as you go feeling you will too as long as you go in with the right mindset. Don't use in with the right mindset. Don't use in with the right mindset. Don't use this to replace fable or judge between this to replace fable or judge between this to replace fable or judge between complex topics. Use this like an if complex topics. Use this like an if complex topics. Use this like an if statement and you'll have a lot of fun statement and you'll have a lot of fun statement and you'll have a lot of fun with it. Let me know how y'all feel. And with it. Let me know how y'all feel. And with it. Let me know how y'all feel. And until next time, peace nerds.

Summary

The main theme is a new AI model called Jev from Typesafe AI, which excels at data classification and returns type-safe JSON. Unlike generative models, Jev is designed for organizing and classifying data with incredible speed and reliability. The practical takeaway is that developers can leverage Jev for efficient data processing, potentially overcoming the slow and unreliable performance issues seen with other AI tools.

View original episode ↗