So, GPT5. 6 has finally been released. In today's video, I'll show you the demos, the levels of intelligence, and everything you need to know to understand just how good this model really is.
So, one of the first things I want to show you all is the intelligence chart. What you're currently looking at is the artificial analysis intelligence index, and this is basically a combined league table for AI models. Now, one thing to understand is that this wasn't previously available before GPT5.
6 was essentially released because these companies didn't have access to the API, and they couldn't really release these benchmarks just yet. Now, currently, what this actually shows us is nine different evaluations, including the GDP Val, Terminal Bench, SciCode, Humanities Last Exam, and other popular, well-respected benchmarks. So, you might be asking, where does GPT5.
6 land? And it's pretty clear. GPT5.
6 lands just right behind Claude Fable 5. Now, I want to make a caveat here because there is something that you should know. Just simply stating that GPT 5.
6 lands behind Claude Fable across a wide range of intelligence tasks doesn't mean that this model is any less capable, or it's not the one that you should be using. And once you understand my reasoning for this, you'll start to understand why. So, the GDP Val is basically just a summary of everything together.
Think of this as your average of averages. So, if you just wanted to quickly look at it and glance where the model is, a rule, this is super valuable. But for those of you guys who are a little bit more, I guess you could say into how the model actually works, and you want to use it on a day-to-day basis, you're probably going to want to understand exactly where it excels and where it underperforms.
And one thing I will say is that I'm just going to put this out there because this is true. GPT5. 6 across most benchmarks is vastly cheaper than Claude Fable 5.
So, even when you see this intelligence analysis index, and you say, "Okay, GPT5. 6 is not on par," do remember that it is more available and vastly more affordable. So, let's get into the next benchmark, which is going to surprise you because this one is called Agent Last Exam.
And this is one where GPT 5. 6 actually outperforms Fable 5. So, Agent Last Exam is essentially a new benchmark testing whether agents can complete actual professional work rather than simply answering difficult questions.
Think of it like this. If you remember the Humanity's Last Exam, that was can the AI answer an extremely difficult expert question. The Agent Last Exam was can the AI take control of a computer and successfully finish an entire expert project.
Now, the agent essentially receives a task, access to a real Windows or Linux computer, relevant files, professional software, and it must independently plan the work, operate the software, create the required deliverables, and leave behind a correct final result. And the tasks include producing animations in Adobe After Effects, building scenes in Unreal Engine, creating engineering models in Siemens NX. And you can see that the pass rate is extraordinarily low.
This spans 55 professional sub-industries, 1,500 collected data sets, and is being expanded. And I don't want to get too too much into this benchmark, but the entire question of this benchmark is can an AI agent generally replace a skilled computer-based worker for an entire piece of work, not just help them with one step. Now, right now that answer is mostly no, but take a look at the agent behavior on GPT 5.
6 So, when we look at the score of this model, and not just the score, but the cost too, we can see that on the Agent Last Exam, GPT 5. 6 so is remarkably better than anything we've seen before. And this goes for a wide range of models.
Now, of course, you could say on some areas Claude Fable 5 is better, but guys, we have to pay attention to what matters here. Look at the API cost in US dollars. The thing doesn't make sense if it's so expensive.
What is the point of using a model where your weekly reset limits are going to be gone after three prompts? Are you going to use that model? I know I'm not.
But a really intelligent and efficient model that allows you to get a lot of work done, now that seems like something I can get behind. And this is where GPT-5. 6-Sol excels.
They definitely did a lot of work to make this model a lot more agentic. So, if you're wondering, where does GPT-5. 6-Sol excel?
This is essentially showing you that on these agentic tasks, it is not only more cost-effective, but more robust in completing those tasks. So, we can see that GPT-5. 6-Sol is pretty remarkable when it comes to that.
So, understand that this is not just a model that is intelligent, it is also cost-effective. Another benchmark I wanted to really take a look at is the Deep SWE. Now, this one, once again, is the specific area where GPT-5.
6-Sol actually excels. Most people will have missed this because they'll look across the board and say, "Oh, Anthropic company, that is the one that is the most agentic, that has the most coding knowledge, and they're just going to far, far exceed. " However, that is not the case in the Deep SWE.
When we look at this, we can see that GPT-5. 6-Sol actually exceeds at 73%, and that makes a lot of difference. You see, Deep SWE is a difficult coding agent benchmark.
Instead of asking a model to write one small function, it gives it a real software repository and expects it to explore unfamiliar codebase, understand com- plicated engineering requests, edit multiple files, run tests and debug failures, and finish a long task autonomously. It contains 113 original tasks across 91 repositories and five programming languages. Every model is placed inside the same mini SWE agent setup, making this comparison one of the best, okay?
And the tasks were written specifically for the benchmark to reduce contamination. So, when you see this result, you can understand that they essentially completed 73% of those difficult questions with a high degree of accuracy. And you might be wondering, why am I not looking at the other coding benchmarks?
Guys, benchmarks are useful, but often times there are mistakes in them. And I will say that even with this benchmark, it does look super good for GPT-5. 6 Sol and the rest of the models.
However, always do your own testing because you have to understand sometimes, whilst benchmarks may hold up, in the real world that may tell a different story. And one last thing I'll add before you move on, is that when we look at GPT-5. 6 Sol, look at the average cost in terms of what it takes to get to that level of intelligence.
You've got $8. 40 versus Claude Fable 5 for $21. 00, I mean $21.
63. That is roughly the same level of intelligence, but for three times cheaper. I know exactly which one I would be picking.
Now, in order to continue this video, what I want to show you guys is some clips from OpenAI where actual real people getting work done with the GPT-5. 6. I'm going to let these videos play because they are super, super interesting.
For example, like this first one where there's a mathematician solving previously unsolvable math problems with the GPT-5. 6. This is one that is absolutely insane.
>> Usually when I get stuck with a math problem, I just want to stop for a while and just get relaxed. So, then I go out to the garden. I just am completely immersed in my own thoughts.
And once I come back, I just run Codex to do all this exploration and all this technical stuff that otherwise it would take me weeks. I remember working on a computer code when I was about five, I think. I barely remember a time I was not able to code.
It was like always in my blood. We've been working for the last three years on a remarkably hard question about algebraic surfaces. We were trying programming stuff, trying pen and paper approach.
We tried previous models, nothing really worked well. I decided, okay, I'm going to try the new Codex 5. 6.
Codex was able to come up for completely new idea. And it helped us disprove the conjecture that we were trying to prove for these 3 years. >> 14 over 5 is actually bigger than 8 over 3.
So, it turns out that the conjecture was basically false. >> Oh, I'm so excited of the discovery. >> [laughter] >> But that's why we do science, right?
Yeah, we we want the fun. >> 5. 6, it it feels kind of natural because you set up the task and the model can recognize that there's like a lot of computation going.
So, it has to spawn the sub-agents automatically without even me asking for this. I see that all these AI tools are meant to empower people. Now, I can focus on what I like doing the most.
So, hard ideas in math, staying with my family, with my kids, and mowing. If you have this kind of audacity to try something very big, you won't be uh scared by the sheer amount of computation because you'll be able to organize it with the model. And I think we are going to have a lot of fun and a lot of new discoveries on the way.
Questions? >> [snorts] >> And we also have an average example where, you know, you can literally start, create, run a business with GPT 5. 6, just showing you that the only limit is your imagination.
>> Build a sell-in timeline tracker with retailer-specific deadlines and contacts for our pumpkin spice launch. >> Three Wishes is a better-for-you breakfast cereal. It is higher in protein, lower in sugar, gluten grain free.
And the inspiration came from our own house and our own family. >> This one I didn't really create, but I thought of the It's called cotton candy. Not trying to glaze myself.
Personally, my favorite. >> We're true team in all the things that we do. >> But there's only so much two people can do.
>> As you grow and the brand becomes bigger, mistakes just have so much more weight. >> In the past, I had so many tech ideas, but I always said I can't do them. I need a technical co-founder.
Codex is that technical co-founder. >> Every time we've had to run our seasonals, they become an operational nightmare. So, we had Codex build this really beautiful command center and dashboard.
>> So, I always start by saying please set a goal for yourself to do this end to end. Please use historical launches as like a means to inform the dashboard itself. Use the Three Wishes branding, especially the new branding.
We have new boxes that we're using. >> It is the most disorganized thought stream that I've ever seen. To them, it's like [laughter] incredibly amazing output.
>> I mean, it wouldn't have been possible before. With 5. 6, I directed it some of Three Wishes brand guidelines and said just make it look like and feel like our brand.
>> So cool. >> Yeah. >> 5.
5 was never capable of doing something along these lines. With 5. 6, I spoke into it for 5 minutes and it output all of these different data slices that they needed.
>> This feels like something you would normally think a big billion-dollar company could build. And you're like, we get this. I have my bespoke software system.
>> Makes you feel real and enterprisy. >> Yeah. >> It frees up the most important time for us.
I get to choose where it goes, whether it's building the brand or spending time with my kids. And that to me is like the the time I'll never get back. And so, it's really exciting.
It It feels like you gave us wings a little bit. >> Another example that I really like is Hiroki, a broccoli farmer running his farm with GPT 5. 6.
Most people don't realize that when you now have access to the world's, you know, essentially smartest intelligence literally at your fingertips, you can do so much more than you thought capable. Essentially, what I'm trying to show you with these open AI clips is that you really want to push for the sky when it comes to asking an AI to do something. I'm not saying you're going to land on the moon, but what I am saying is that things that used to be out of reach for the average person simply aren't anymore.
Intelligence at your fingertips isn't something to be wasted, it's something that you should be curious with, and it's something that you should explore. Maybe there's something out there that you can do that you simply haven't utilized yet. And in this example, that's where we see Hiroki use this to run his farm more efficiently with GPT-5.
6. >> I'm in a broccoli field right now. And it's hard to tell where I've already worked.
So, I want to build a system to map out those boundaries digitally. When I first came to Hokkaido, it was really tough because I didn't know how to drive a tractor or grow vegetables at all. I'm not an engineer, but I'm using Codex to make things more efficient.
I used to manually roll up the vinyl on the sides of the greenhouse for ventilation. But with Codex, I was able to automate that process by installing electric motors. I built a system where I can control these electric motors going up and down from anywhere as long as I have an internet connection.
With the old model, it was a constant loop of back and forth, making tweaks and giving instructions. But with the new model, Codex reads my database from a single prompt and uses various tools on its own. Even without me constantly checking or managing it, it figures out how to reach the final goal and executes it.
It's just an unbelievable change. Being able to build a system myself that actually works in the field, it gave me the sense of wonder. Like, "Wow, what is this?
" It felt like I was watching magic happen. >> [laughter] >> That crowing is so loud. We can't get through the interview.
How do you make a chicken quiet down? >> [laughter] >> Now that you've seen that clip of Haruki, let's actually get on to some of the things I found on Twitter that were actually community use cases that I thought were really cool. So, one of the ones here was essentially this two-time speed video of this computer using agent.
And this is GPT Soul. Now, one of the things you'll realize is that this is moving absurdly fast. And I think most people will miss the significance of this event.
And the reason I say that is because most people do not realize that what currently we have intelligence that is correct, but we don't have super fast intelligence. What do I mean by that? Well, most people don't realize that the current AI models that we have, the way they talk to us, it's incredibly slow compared to what is currently available out there.
The current speed that you're seeing is not the standard version of GPT 5. 6. It is a special version.
This is actually called fast mode. And anyone with GPT 5. 6 Ultra on fast mode can essentially do this.
This is no MCP. This is essentially just running this at higher tokens per second. Now, of course, this does mean you run through your usage super quickly.
But imagine having an AI that could run, I don't know, let's say 10 times as fast as it is now, but you have access to that unfettered, untapped intelligence, okay? Like imagine asking your AI to code your website, okay, a complete website, something super complex, and it does it within 30 seconds. That is what we are talking about here when we look at the tokens per second being increased.
And later on in the year, there will be Cerebrus with 750 tokens per second. And this is going to radically change the game and usher in a new area of intelligence. And I think that is when AI is going to become really, really fun because agents are going to move like 5 to 10 times as fast and we're going to see some incredible things.
Let's take a look at this example from this developer who used GPT-5. 6 to essentially create a working version of Google Earth. I wish I was as creative as he is to come up with this idea.
Of course, I'll leave a link down to this stuff so you can go ahead and check out its page. I will leave a links to all of this. But this is a really, really cool idea.
You can see that this application works. You can see that it actually has a fly-through called icons of the world going through all of the different, you know, icons around the world, the Big Ben, the Eiffel Tower, and it's able to essentially lot different things such as buildings, airports, mountain peaks, rivers. I mean, this kind of thing could be used for, I mean, so many different use cases.
I don't even want to think about the amount of different use cases we now have when we have such a more capable model. And you guys might think I'm exaggerating, but I've tested this model myself and I'm going to show you guys some examples as well of just how good this model is when it comes to coding things on the fly. Remember guys, this is a model that is essentially at the top tier of intelligence and at the top tier in terms of the price range.
So you're essentially getting the best of both worlds. And this is ushering in a new area of creativity because whilst with Anthropic models, you're usually limited by your usage, with OpenAI, you are able to basically do what you want. Once again, we had another example from Efron Mullick.
This is where he wanted to build a procedurally growing city. And this is something that once again seems very, very impressive. The ability to code this entire thing from scratch with being procedurally generated.
Essentially just I think it's coded in either HTML, 3JS, honestly I'm not entirely sure, but we have moved so far from where we were just a year ago and I can't imagine where these models are going to be in a year or two years from now. You're going to be able to literally build anything at the click of a button and the only limitations are going to be your use cases. Maybe you'd like to, you know, plan some developments in your own city, maybe you'd like to plan something in your living room, I don't know, work on your own creative engineering project.
The ideas and the thoughts are completely limitless. This is something that I just think is absolutely insane once you start to get your use cases down. And this was one of the most impressive examples that was floating around on Twitter where GPT-5.
6 Sol one-shotted this voxel-based Manhattan. There is incredible precision with this build and apparently this ran for a week. So this is Matt Shumer and he essentially asked GPT-5.
6 Sol to create a voxel-based Manhattan. Now I believe he was using the goal mode in GPT-5. 6 and if you don't know what that is, that is a mode where you essentially ask the AI to go out and achieve a goal and it literally will not stop until that one specific goal is done.
And in this case, this allegedly took an entire week to create, but with the level of detail, I can completely understand why. And like I said guys, this is just the beginning. Big and frontier models are coming fairly soon and you're going to be able to do even crazier projects.
Now OpenAI have of course introduced some of their own use cases, some of their own things that were built with GPT-5. 6 Sol. And in a moment, I'll show you guys exactly what those look like and how you can access them because they help you in terms of your creativity if you aren't as creative as some of the people you're currently seeing here.
So on the OpenAI page, one of the things that I think is really cool is that they come across this stuff right here which is essentially showing the leap forward in design. And it talks about the fact that the front-end capabilities turn natural language requests into polished, interactive explanations and visualizations within ChatGPT. So, this is something that I think is really cool.
So, for example here, you can see that this is an interactive spirograph explaining how it works. This is something that I think is really cool. And this is something that you need to familiarize yourself with because this is the concept of generative UI.
Now, if you're wondering what generative UI is, it's user interfaces that is essentially developed on the fly specifically for you. And I think this is going to become increasingly more commonplace as intelligence becomes more and more available. So, let's say you want to learn something in a super specific style.
Sometimes, what you can do is ask an AI to code a specific visualization that you can change with and it codes that on the fly, a small application that builds you, and then you can use that to understand complex concepts. Another one here is the interactive wave inference. We can see this entire diagram is super super useful for people who are trying to understand this concept.
And another one is the interactive GPT tokenizer. So, when using GPT 5. 6 Soul, this is something most people are going to miss.
They don't realize that Google and OpenAI are pushing forward in this front of interactive generative UI. So, here you can see you're able to type in the text and then it shows you how it is being tokenized. So, I think this is something that is super super useful and we will start to see more and more of this, like I said, as the intelligence becomes even better.
So, definitely don't forget to try out this prompt, can you interactively explain how XYZ works? Because this is going to become something that is more commonplace and you'll find that it is often better than standard explanations, especially for those of you that are visual learners. Another thing we can see is that they say that GPT 5.
6 delivers a step change in design judgment. With only high-level direction, GPT 5. 6 creates tasteful, ergonomic, and functional interfaces.
It has stronger compute use capabilities that it inspects and refine the rendered result, not just generate the code. So, it can essentially change those finishing touches. And here is where we get five different demos to show us just how great that is.
Number one is the salt wind, which is essentially where you read the wind, thread the gates, and chase a faster line across the sunlit sea. So, this is a game completely coded by GPT 5. 6 salt.
And remember guys, this is stuff that you can go ahead and code yourself. I'm not sure exactly what kind of things you'd want to go and code yourself, but you have to understand the reason they're trying to showcase these capabilities is because there are so many things that have to work in order for a game to not have bugs, not to have consistent crashes, and just to function very effectively. This isn't the only one.
There was this tiny voids game. Here you can see this is the game that is pretty local, a tiny voids game that allows you to essentially suck in cubes. I mean, this is once again like not like an indie game, but this is something that I guess you could call something that is really fun depending on the kind of things that you want to do.
And like I said guys, of course this just does depend on your creativity. What I will say is that this is where you get to see that creativity is your only limit. And when you start to realize that, you start to realize that okay, maybe I can do a lot more with this thing than I initially thought.
Maybe I could code a visualization for this or that. Maybe I could code a web app for me to play a specific game that's going to help me understand more concepts. These are the kind of things that make you realize just how crazy AI is.
This wasn't an entire company that built this entire game. This was GPT 5. 6 on one prompt.
And so, when we actually take a look at the websites here, something we can take a look at is the front end design actually having taste. For the longest time we're being stuck with those AI generated websites and having to use all of those janky ways to figure out how to get the websites to actually look good. But now, natively GPT 5.
6 because it can actually look at those designs and come back with something else, you start to realize that it actually understands exactly what it's generating and so therefore can generate something a lot more coherent. And this is something that I think, you know, when you're using the model, it's something that's subtle, but it's something that's useful because, for example, presentations like this, you'll just start to realize that, okay, maybe I don't need to go ahead and do three reviews. Maybe I just need to do one review before approving.
So, this is something that I think is really cool. Now, one of the big questions you're probably going to have is, okay, I understand what GPT-5. 6 saw there is, but really, what is it that makes it so good?
And is it better than Fable 5? So, essentially, what I did was I gathered some of the opinions from people that had been testing this model for months. I could give you guys my opinion from testing it for the last, you know, 12 hours, but to be honest, that's not enough depth for the people that have been testing this early that have had exclusive OpenAI access.
And them having it for 2 months certainly should say something. So, this guy, Mashima, essentially has had GPT-5. 6 since May 27th, okay?
And for the first 2 weeks, he said that this model was super impressive, but the TLDR is that Claude Fable is better than GPT-5. 6 in certain ways. The benchmarks may look close, but apparently Fable has a big model smell.
GPT-5. 6 feels like a smaller model that has had reinforcement learning making it appear very well, and the difference that shows up the moment you push past normal coding work. And apparently, it shows up in the trust, too.
Now, he does say that GPT-5. 6 to do ambitious work, you may need to steer it a bit, but with Fable, you can simply just describe the final destination once, and it gets there completely autonomously. Now, he does add that GPT-5.
6 does beat Fable in a few places that matter, you know, the interface, the willingness to do security work. It's now his security auditor and his second to advise, but other than that, he does take Fable. Another example here from Nate Hurk actually says that Fable is the manager and Sol can be the worker.
And there are many different cases where Fable refuses to answer and Sol can step in quickly and generate a very, very effective answer. And so, this is something that I want you guys to understand. Both of them essentially are pointing toward the fact that GPT-5.
6 Sol is very good and is very token efficient, but it feels like a less capable model. Now, I'm not an open AI fanboy, but I do want to basically make this case. Why don't people remember the fact that model releases are essentially a part of a model family?
And if we look at the model families, we can see that there are completely two different model families, and I think people are making a mistake. So, when you look at the top row, you see the Opus 4. 6, it's followed by Opus 4.
7, and it's followed by Opus 4. 8. And then, what do we have?
We have a big model jump towards Fable 5. Now, look at GPT series. GPT 5.
4, GPT 5. 5, GPT 5. 6.
And why are we now comparing that 5. 6 model to Fable 5, which is, as you would argue, the next complete step change in terms of intelligence and raw ability? So, I would argue that GPT 5.
6, by all means, should probably be compared to Opus 4. 8, because I'm not sure Open AI decided to train an entirely new model. They seem to be sitting on this one for months.
So, the real test for Open AI should be looking at GPT 6 versus Fable 5, and in my opinion, this seems like a bit of an unfair comparison. So, when you look at Opus 4. 8 and and GPT 5.
6, then the picture starts to change. Comparing those tiers of models, GPT 5. 6 completely wipes the floor with Opus 4.
8. With that being said, I think GPT 5. 6 is very good for what most people are going to get.
Most people will never need the intelligence that Fable offers, but either ways, I'd love to know your thoughts and what you've been using this for. If you enjoyed the video, I'll see you in the next one. Now, one last update I would love to talk about is something that most people will skip because it isn't on the front page, and it isn't some insane benchmark.
This is Chat GPT for work. So, this thing that they released, I think, is probably going to have the most impact in terms of actual usability. Essentially, it's an agent in Chat GPT that helps you take on more ambitious tasks across your apps and workflows to generate finished materials like slides, documents, web apps, and stay with complex projects for hours by breaking them into smaller steps and completing them independently.
So, essentially what this is is this is basically Codex but for work or as some of you may realize, it's kind of like Claude co-work but for ChatGPT. And so, this is something that was essentially a coding agent for developers. A lot of people now use it for work outside of software development.
And what I will do as well as if you want to get started with this, I'll leave a link to my Codex tutorial in which I show you how to actually use Codex for everyday work because a lot of people would see Codex, say I'm not technical, I'm not use it. But trust me, it was super easy to use. Of course, I'll have an updated tutorial coming on this.
But essentially, ChatGPT for work is just an optimized version of GPT 5. 6 with Soul embedded in it. And I would argue that if you're going to do this, download the Codex application because essentially what happens is is that when you are on the desktop version, that is where you will get the most essentially comprehensive experience of ChatGPT for work because you can work with the files that are on your computer and it just leads to a lot more efficiency in terms of being able to work that fluidly.
So, with that being said, I think most people don't realize just how crazy ChatGPT has gotten. And I will say, don't forget to check out the live voice because it is exceptional. And most people aren't even going to test it.
So, with that being said, if you guys have enjoyed today's video, hopefully this has been a super duper good one. And if you have any questions about GPT 65. Soul, I'd love to know your thoughts and feelings.