Which AI Should You Actually Use? Claude vs ChatGPT vs Grok vs Gemini (2026 Guide)
Josh:
Over the last 90 days, Frontier Labs shipped 15 plus models.
Josh:
OpenAI shipped three, Anthropic shipped four, Google shipped three,
Josh:
even Meta and Elon Musk shipped Frontier models. The Chinese shipped a bunch as well.
Josh:
And so if you're listening to the show, you're probably wondering,
Josh:
which model should I be using right now?
Josh:
Most of you likely have a GPT or Claude subscription, but you're wondering,
Josh:
should I be using these different models?
Josh:
The truth is, the Frontier landscape of AI models, when it was originally thought
Josh:
to be one or two, has now expanded to hundreds and hundreds of models.
Josh:
In OpenRata alone, you can access 400 plus. And so on this episode,
Josh:
we're going to dig into which model you should use for what particular use and
Josh:
when makes the most sense.
Josh:
And you'll realize that the argument has shifted not from using the most intelligent
Josh:
model, but maybe using the most cheaper model or using the model that's specific for you.
Ejaaz:
You got options, baby. There's a lot going on in the AI space and we're going
Ejaaz:
to help navigate that space because there is a lot of options and it does get
Ejaaz:
overwhelming. I think between the two of us, we've probably touched every single
Ejaaz:
one of these, at least a couple of times.
Josh:
I have too many subscriptions, Josh.
Ejaaz:
So yeah, between our 75 different subscriptions, we have covered these.
Ejaaz:
We have some feedback about which one to use for when, which is best for which use cases.
Ejaaz:
And I guess to start, we have this like pretty helpful visual companion artifact
Ejaaz:
here that can walk through kind
Ejaaz:
of what we're thinking, how we think about this. The first is the split.
Ejaaz:
There are two distinct classes of model when it comes to considering which ones
Ejaaz:
to use. The first is open source.
Ejaaz:
These are Chinese models predominantly. Most of these are your Kimis,
Ejaaz:
your DeepSeaks, like all of those models are Chinese open models.
Ejaaz:
And then we have the actual US models that are all closed source.
Ejaaz:
There's a big difference between the two. And we could see if we scroll down
Ejaaz:
a little bit, the difference in usage between these two, because in June of
Ejaaz:
2025, US Labs accounted for 70% of the tokens generated. And now they're down to 30%.
Ejaaz:
The economics have changed widely. So now that 30% is worth much more than 70%.
Ejaaz:
But it is interesting to note that there has been this kind of reinvigoration
Ejaaz:
of Chinese models over the past couple of months. Now, granted,
Ejaaz:
this is based on open router data.
Ejaaz:
Open router is a single model router instance. This is not reflective of the norm.
Ejaaz:
But it's just worth noting that some of these open source models are pretty
Ejaaz:
powerful for the task at hand.
Josh:
Yeah, I think no one can debate the fact that people have shifted a lot using
Josh:
these open source models. And it's for a variety of different reasons.
Josh:
People want to own their own data, they want to run it privately at home.
Josh:
But the biggest shift, the biggest reason has been because these models are
Josh:
a lot cheaper. And open router data you mentioned only really represents a very
Josh:
niche sector of software engineers that want to experiment with a bunch of these models.
Josh:
But even in the enterprise world, where OpenAI and Anthropic are pretty dominant
Josh:
with their own model share, they've been losing the token market share to enterprises
Josh:
that are trialing and testing different models to save tens to hundreds of millions of dollars as well.
Josh:
And when I look at the main reason why, it's not because the open models are
Josh:
the most intelligent. You'll see in a second as we go through the scorecard,
Josh:
GLM 5.3 from China, amazing model, not as good as Fable 5. You'll look at Kimi
Josh:
K 2.7, you'll think the same thing.
Josh:
Those models are good enough to do the bulk of your own work.
Josh:
It doesn't sound too much like AI slop.
Josh:
It actually just speaks to you normally. And the biggest advantage is it has
Josh:
fewer safeguards, which has its pros and cons, but it basically does the tasks that you ask it to.
Josh:
Now, overall, it's fine if we talk about a bunch of models on this show,
Josh:
but it's good to get an overview, essentially, of how intelligent or how effective these models are.
Josh:
And what we have on the screen here is something known as the Artificial Analysis
Josh:
Index or Intelligence Index.
Josh:
And this is basically the best benchmark to test the general intelligence of
Josh:
these different models.
Josh:
Now, it may come as no surprise to you, but the Claude models,
Josh:
Anthropix models, top this. We've got Opus 5 at the top, which is their most
Josh:
recent model launch. We've got Claude Fable 5.
Josh:
And then, surprising to me, Josh, but Grok 4.6 from Elon, it is just,
Josh:
it is fantastic how effectively SpaceX has been able to pivot from being way behind the model race.
Josh:
They fired their entire AI team, hired a bunch of new folk, acquired Cursor
Josh:
for $60 billion and put out a model in, I think it's like the last six weeks
Josh:
that was able to contend with the top. And the best part is he has all the GPUs
Josh:
and compute to train them.
Josh:
And then if you go further down the stack, you'll see the likes of OpenAI,
Josh:
GPT 5.6 Sol, and a bunch of other Chinese open source models.
Ejaaz:
The question is, what do you make of this? As someone who is just a user of
Ejaaz:
AI, what do you make of this? How do you navigate this?
Josh:
Let me trial this, Josh. Let me ask you this question. Of these models that
Josh:
you see on the screen here, which ones are you using and for what specifically?
Ejaaz:
It's funny because about 90% of the usage comes from the top two.
Ejaaz:
And then 10% comes from everything else. And this is kind of like what I was
Ejaaz:
going for. I'm actually curious, is it a similar thing for you?
Ejaaz:
Like what kind of models are you using?
Josh:
Yes, but okay. So I would say it's around, I'm going to say like 75, 25.
Josh:
And the reason why I have that slight adjustment is because I kind of use the
Josh:
AI model that's most convenient to me wherever I'm getting the information or
Josh:
the inclination to use an AI model. So X is a perfect example, right? I'm scrolling X.
Josh:
I read something from one of these genius AI researchers and I'm like,
Josh:
I have no idea what the hell this means.
Josh:
I tap the Ask Rock button. I'm on Google. I'm searching something, right? I'm on my laptop.
Josh:
I'm like, okay, I'll just speak to Gemini in this. So in that sense,
Josh:
wherever the AI is conveniently placed, I use it there.
Josh:
And I have a feeling when Apple releases like their new AI model in a couple
Josh:
of months or a couple of weeks, actually, I'll probably use it on my phone as
Josh:
well because the models just generally are good enough to answer my basic questions.
Ejaaz:
Yeah, okay. So we're looking at these charts and we're seeing,
Ejaaz:
okay, here's where the intelligence is. Here's where things get a little less
Ejaaz:
intelligent, but much cheaper.
Ejaaz:
And if you're a normal consumer, I have a feeling that cost isn't that big of
Ejaaz:
a deal because the cost differences are really at scale.
Ejaaz:
What a lot of these companies offer, what Anthropoc offers, what OpenAI offers,
Ejaaz:
what Grok and Gemini offer is just the $20 a month, $100 a month, $200 a month plan.
Ejaaz:
You could just pay a fixed amount and get access to a lot of usage of these models.
Ejaaz:
So if you are a general consumer who is able to spend $20 per month on a subscription,
Ejaaz:
you're probably better served going to one of the major iolabs.
Ejaaz:
The product is better. You're going to get more intelligence.
Ejaaz:
You're going to get a more coherent product.
Ejaaz:
I find that a lot of the people who I speak to a lot of the time in my experience,
Ejaaz:
the only real reason to use these open source Chinese models is if you're consuming
Ejaaz:
tremendously large amounts of tokens for, say, agentic tasks.
Ejaaz:
Like if you're running an open claw instance or any sort of claw with a lot
Ejaaz:
of automated workflows, a lot of agents, you may want to defer to some of these
Ejaaz:
cheaper, perhaps the Chinese open source models just to save on costs on the low end.
Ejaaz:
But when it comes to day to day use on the high end, I'd say probably 90 plus
Ejaaz:
percent of the people watching this are best served just getting a subscription
Ejaaz:
to your favorite ai company
Ejaaz:
Using that product. That's kind of how I use it. It's funny that the remaining
Ejaaz:
10% of prompts that aren't done through Anthropic are done mostly through Grok.
Ejaaz:
Like you mentioned, when I'm on X 24 hours a day, it's a good companion because
Ejaaz:
it has access to that database.
Ejaaz:
And then the other one is actually ChashyBT because they have a really amazing
Ejaaz:
image generation model.
Ejaaz:
And I love creating memes or sending images to my friends who are just like
Ejaaz:
using that as the visual component. I find that really compelling.
Ejaaz:
So that's kind of my stack is like mostly Anthropic, sometimes Grok,
Ejaaz:
sometimes chat gpt particularly for the image gen i think when it comes to image
Ejaaz:
gen we used to use nano banana and gemini all the time that was like the top
Ejaaz:
dog i think ever since image gen 2.0 release from chat gpt
Ejaaz:
is very compelling what i like is that it does get actual thinking behind it
Ejaaz:
so previously when you asked to generate an image it would take your words at
Ejaaz:
face value generate the image now it actually does some inference it does some thinking
Ejaaz:
it kind of interprets your message and interprets what it's seeing in the image
Ejaaz:
and then creates a much better output. So that's kind of how I would think
Ejaaz:
of navigating this um like chinese models are certainly a big deal but unless
Ejaaz:
you are kind of advanced super user consuming a ton of tokens perhaps they're not that valuable
Josh:
I will speak from my own personal experience i
Josh:
use the clawed models pretty much all the time
Josh:
but i've started using opus 5 more than fable 5
Josh:
and and this might sound like a silly reason but it's because it speaks to me
Josh:
like a normal human being i don't know what the personality dials are on opus 5 versus fable 5 but
Josh:
they need to do that to fable 5 pronto because fable 5 just waffles and speaks
Josh:
in like archaic prose and i'm just like listen just give me
Josh:
give me the information that i've asked for and do not give me anything else
Josh:
i don't want to have to read,
Josh:
an essay every time i use this right now grok 4.6 i've actually noticed the
Josh:
intelligently josh i don't know if you've tried this on x you probably do when
Josh:
you're like kind of like trying to figure It speaks to you super intelligently
Josh:
and there's less crass about it.
Josh:
If you remember the earlier versions of Grok used to kind of like throw in some
Josh:
kind of like questionable words, slang and phrases.
Josh:
And now it just speaks to you like an intelligent human being.
Josh:
And I think this is because Elon's like, you know, SpaceX is IPO'd.
Josh:
He's got Cursor, very professional researchers involved, and they're trying
Josh:
to target a more enterprise oriented audience. And he said before that Grok
Josh:
4.7 is on its way to being released in a few weeks now. He's aiming for Grok
Josh:
5 by the end of the year. That's literally only in a couple of months.
Josh:
He is really targeting the enterprise landscape because he's seen Anthropic and OpenAI go after it.
Josh:
Now, same as you, on the Chinese models, I am not using any of them.
Josh:
And I admit this might be my own bias because...
Josh:
Mainly, I'm like, I'm not sure, like, I want to go to the efforts of signing
Josh:
up on an account, using those models, and then being like, well,
Josh:
I don't know whether this is giving me the right information.
Josh:
I don't know. I just kind of have this weird trust thing with the US and American
Josh:
brands. And maybe that's like my own fault.
Josh:
And then the one thing that I'll point out across all these different models
Josh:
that we make the point on on the screen here is.
Ejaaz:
I think a bulk
Josh:
Of the companies, Anthropic and OpenAI especially, have focused on making their
Josh:
models really good for enterprises. But as I mentioned earlier,
Josh:
the shift has really happened to open source because...
Josh:
More available models are cheaper and more accessible to people to use. They can run privately.
Josh:
But Anthropic and OpenAI's response to this is, we'll just distill our main
Josh:
model and give you cheaper models.
Josh:
So that's what OpenAI has done, right? You've got GPT 5.6. There's three versions
Josh:
of it. You've got Sol, Terra, and Luna. Luna is, I think, slashed 80% of the
Josh:
price than it was before already a week ago, right?
Josh:
So now it's like competing very much with GLM 5.3.
Josh:
And you've got Anthropic doing the same. I think Opus 5 is like half the cost
Josh:
of Fable 5, whilst being as capable as Fable 5.
Ejaaz:
And they also kept the SANA pricing low too, which is the lowest layer.
Ejaaz:
So that's like, yeah, there's plenty of options there.
Josh:
Plenty of options. And I think this comes down to one thing,
Josh:
and I don't want to make this about compute, but I have to, Josh.
Josh:
I think whichever lab, whether you're Chinese, whether you're like a low tier
Josh:
US American lab, or whether you're the high tier labs, compute will dictate
Josh:
whether you can serve these bottles cheaper, which will basically decide or
Josh:
determine which customers use your product.
Ejaaz:
Yeah. And I mean, this is ultimately coming down to your use case.
Ejaaz:
Like a lot of people won't ever run into this problem because they will never
Ejaaz:
need this thing called an API key.
Ejaaz:
They'll never actually pay per token. They'll oftentimes just pay through the subscription.
Ejaaz:
And if you're using a subscription, you have a couple of options.
Ejaaz:
Maybe this is a good time to go through these subscriptions and what are the
Ejaaz:
offerings of everybody. We have Gemini, which perhaps we could start there because
Ejaaz:
I feel like we've been mean to Gemini recently. We haven't talked much about Google.
Ejaaz:
What's that one? Yeah, like that one.
Ejaaz:
Who's that? They're currently at Gemini 3.1 Pro, which is nowhere really near the frontier
Ejaaz:
um anymore unfortunately they do have nano banana pro which is like this fun cool image generation
Ejaaz:
process they have what i find interesting about google and if you are interested
Ejaaz:
in these tools it may be interesting for a subscription google has these like
Ejaaz:
weird edge case tools that have really interesting harnesses one of them is for music production
Ejaaz:
i remember we've used this on the show many times it's really good at generating
Ejaaz:
lyrics and generating music in a way that sounds very good
Ejaaz:
I've only had that experience on a Google product. I haven't been able to use that anywhere else.
Josh:
Lyria.
Ejaaz:
Lyria, that's right. Yes, Lyria, Project Lyria.
Josh:
Remember we created the jingle for the show?
Ejaaz:
Yeah, so good, dude. So like Gemini and Google is like good for the weird stuff.
Ejaaz:
Like it'll generate you pretty good images.
Ejaaz:
It'll generate you fun music. They have this other tool.
Ejaaz:
This is in like the Google Lab suite where it'll generate you marketing material.
Ejaaz:
So if you have a lot of material from your brand or you have a bunch of logos
Ejaaz:
and you want marketing material, kind of say your logo's printed on something
Ejaaz:
that's staged nicely or you want custom merchandise, It'll do a lot of that
Ejaaz:
for you. So it's fun for those kind of narrow use cases of where the lab's products lie.
Ejaaz:
Outside of that, I don't know, I can't really recommend it that much.
Ejaaz:
It's not the strongest, not close to the strongest.
Ejaaz:
I'd say probably on top of that, we have Grok, which is very quickly catching up.
Ejaaz:
If you spend a ton of time on X like we do, Grok is your best friend.
Ejaaz:
Grok has access to that data that is available in real time on X.
Ejaaz:
If you're looking for current up-to-date news on pretty much anything, Grok is your go-to.
Ejaaz:
It can sort and cite specific tweets from specific moments that are happening
Ejaaz:
in real time and just generally speaking like you mentioned details it's very
Ejaaz:
good at just communicating with you directly
Ejaaz:
there's not a lot of fluff it is a science and technical model it is direct to the point
Ejaaz:
it'll get vulgar with you if you want it's fun to play around with there is
Ejaaz:
the voice mode which is really fun that i've i'd say of all the voice modes
Ejaaz:
that's the voice mode i've had the most fun with is because it's just so
Ejaaz:
ridiculous and oftentimes if you see memes online it is because is the voice behind it.
Josh:
The GPT voice mode just jars me so much. I don't know, have you spoken to it.
Ejaaz:
With the recent update? The new version, I have, yes. So when you say that, what do you mean?
Josh:
And so it sounds more human, right? It like pauses, but it sounds like it's
Josh:
like some sassy person that's talking to me. It does. Like I asked this question and then he goes,
Josh:
hmm, yeah, yeah, like, yeah, like, listen, like, I get it.
Josh:
And, you know, this is what the weather is. And I'm like, yo,
Josh:
I don't need this. I just asked you what the weather is.
Josh:
Like, you know, just tell me what's up. But I agree with you largely on like on Google and SpaceX.
Josh:
What I will say about Google, Josh, is I think they have a similar profile to
Josh:
Microsoft. I'm interested if you agree with me here.
Josh:
Microsoft is not a name that I would draw against the top frontier AI Model
Josh:
Labs. In fact, I think they fumbled the bag massively.
Josh:
But they are embedded across pretty much every single Fortune 500 company,
Josh:
whether we like it or not. Like we just live in our tech bubble.
Josh:
But outside of the tech bubble, people run Microsoft Teams and the Microsoft Suite.
Josh:
And I think Google's in a similar position, essentially, like a lot of new upstarts
Josh:
use Google Docs, Google Suite, and they're just going to click the Gemini button
Josh:
in the same way that I do so because it's just convenient. Do you agree with that?
Ejaaz:
Yeah, I guess there's just a moat to being like the biggest company in the world
Ejaaz:
and having access to all of these kind of legacy enterprise companies.
Ejaaz:
It feels like anyone who has built a business over the last 20 or 30 years has
Ejaaz:
been using Microsoft, has been using Google.
Ejaaz:
They're going to continue to leverage that. It seems like Google is slightly
Ejaaz:
more ahead than Microsoft, if I had to guess. I'm much more excited about Google
Ejaaz:
as a company than Microsoft is, but they both benefit from that legacy customer base.
Josh:
Well, both are amazing VCs, right?
Ejaaz:
Both are incredible VCs. Yes. I think they should be judged not on the quality
Ejaaz:
of their product, but the quality of their investment in the other competitors
Ejaaz:
that are going to crush them.
Ejaaz:
But I do like that, right? It's like they have a hedge. Microsoft owns a large
Ejaaz:
part of OpenAI. That's incredible.
Ejaaz:
Google owns this like massive stake in SpaceX, in Anthropic,
Ejaaz:
in both of these companies that are going to like be huge IPOs,
Ejaaz:
one of which already did.
Ejaaz:
So they are amazing venture funds, perhaps slightly less better AI labs.
Ejaaz:
The two that don't have to benefit from this kind of legacy software are the
Ejaaz:
two that we're probably going to recommend.
Ejaaz:
And in fact, if you scroll down a little bit further, we could see the box of
Ejaaz:
price against intelligence and kind of where each of these models sit.
Ejaaz:
Yeah, this chart right here.
Ejaaz:
And we're looking at those in the top right. this is where most of the people
Ejaaz:
are going to want to live this is where i think we spend all of our time you're
Ejaaz:
either getting a chat gpt or again a cloud membership and like that's that's
Ejaaz:
for the most part that's what most people need uh the differences are
Ejaaz:
small but noteworthy mostly as it relates to the models i mean both of them
Ejaaz:
i can't recommend enough like spend the 20 bucks a month try it just try it
Ejaaz:
for a month if you don't have it see what you think if you run out of tokens
Ejaaz:
upgrade to 100 a month this is access to the same intelligence that all these
Ejaaz:
researchers are making math breakthroughs with
Ejaaz:
and it's really impressive and really capable now each
Ejaaz:
one of these is slightly different i'd say claude is a little more tailored towards coding
Ejaaz:
writing and it's a really strong general purpose model i find with gpt 5.6 soul
Ejaaz:
in particular it's very good at being the
Ejaaz:
High level operator so if you're working on like complicated tasks it's good
Ejaaz:
at creating spec sheets it's good at running checks against the things that
Ejaaz:
you do and then oftentimes like i find that that claude and the fabled models
Ejaaz:
and the opus models are very good at implementation
Ejaaz:
and they're very good with less guidance so the way i've been using these models
Ejaaz:
recently is with less and less guidance i think earlier on
Ejaaz:
i had this whole skill sheet and it was like 15 different skills that would
Ejaaz:
automate different things for me and slowly over time that skill sheet has gone
Ejaaz:
lower and lower and lower and we were talking actually before the show
Ejaaz:
the best way to extract the most value out of these models is just to
Ejaaz:
get out of the way give it the end goal and say like hey go do this thing for
Ejaaz:
me i trust that you have the intelligence to do this i'm not going to put guardrails
Ejaaz:
on it you have the context you can go and do this for me and i think that's
Ejaaz:
mostly where i find myself using this personally is i use fable
Ejaaz:
i i do not leave any fable tokens unused very valuable tokens each week and
Ejaaz:
then opus is the fallback for the general workhorse model and it's been a really
Ejaaz:
powerful combo of just kind of
Ejaaz:
being able to accomplish anything that you want it connects to all of my
Ejaaz:
services i have it connected to my email to my google drive it has access to
Ejaaz:
my files the folders the context
Ejaaz:
and it's just a really helpful all-in assistant it works really well.
Josh:
I think people are also probably wondering, okay, well, what if I don't really
Josh:
care about the general intelligence as much?
Josh:
What if I'm trying to do a specific thing? Which model should I use then?
Josh:
I think there are two things to consider here. Number one is,
Josh:
if you are a software engineer, like AI models have been all the rage for their
Josh:
coding capabilities, but maybe a bunch of you actually like code hardcore companies
Josh:
and you want to figure out which model you should use.
Josh:
I would say in that question, the clawed models and the GPT 516 models are pretty
Josh:
high up there. Codex usage has gone from 5 million users to 15 million users
Josh:
in about a month and a half.
Josh:
I am tired of seeing Thibaut, who is the head of Codex at OpenAI,
Josh:
tweets on my timeline every single time he gets a million user update.
Josh:
But these models are very powerful at not only understanding and reading your
Josh:
code base, but intuitively figuring out what product or feature you should build next.
Josh:
I have a ton of feedback from friends that work at companies that are tech adjacent
Josh:
and maybe not even in tech at all, which use these models to build their premium
Josh:
features. The other thing I'll say is.
Josh:
You know, we're talking about the top right box over here, which is essentially
Josh:
the Pareto Frontier. So if you had to sacrifice something of cost or intelligence
Josh:
or blah, blah, blah, you would still be using these types of models.
Josh:
But if I had to take a bet, if I was a betting man, I would say Grok 4.6.
Josh:
Quen 3.8, and Kimi will be inside this box within a couple of months time.
Josh:
That's going to change the way we use these different types of models.
Josh:
I said before we started this show, it's unsexy to say.
Josh:
But I think the number one AI company that will come out in the next 12 months
Josh:
will be some form of aggregator platform. We saw that Stripe just acquired Open
Josh:
Router for $7 billion, allegedly.
Josh:
We put out an episode of this yesterday. Definitely go check that out.
Josh:
But I think we're going to see these platforms that help you pick and choose
Josh:
which models to use at the right time.
Josh:
It aggregates your memory so it already knows what you want to do.
Josh:
And so you don't have any of the complications around that. I think we'll see
Josh:
a bunch of these models kind of step up there.
Josh:
Now, aside from coding, if you aren't a coder, but if you are,
Josh:
let's say, a general knowledge worker, you go to work, you use email,
Josh:
you use Slack, you use a bunch of these other plugins that you also mentioned.
Josh:
There are models that are specifically good for agentic tool use.
Josh:
And actually, my most recent favorite is Grok 4.6 or GrokBot specifically.
Josh:
This was released from Elon Musk and SpaceX, I think last week.
Josh:
And it basically is an agent or multiple agents that spins up in a virtual machine
Josh:
on like in the cloud. So you don't have to worry about it running on your laptop
Josh:
and hacking all your stuff.
Josh:
And you can give it access to any and every tool and it intuitively understands
Josh:
and gets what you want to do.
Josh:
It can learn what you do over time and it improves.
Josh:
So I think when you look at the Metamuse Spark 1.2 and you might think,
Josh:
oh, that's a meta model. I don't ever use meta.
Josh:
These types of models are gonna become more available for general usage or knowledge work.
Josh:
And I think if that's something that you're inclined to use and you don't really
Josh:
care about the Google search stuff, you'll use Claude for that anyway.
Josh:
These are models that I'll probably look into because they are superior in many ways.
Ejaaz:
Yeah, it's fun to pick one and stick with it. I find that like oftentimes with
Ejaaz:
the aggregators, the right time to use that if you're doing lots of work,
Ejaaz:
if you're generating lots of tokens, if you are just a person who wants to use
Ejaaz:
AI for their day-to-day tasks to help you come up with like a grocery list, to help you
Ejaaz:
cook specific things, to help you go to the gym and give you workout classes and ideas.
Ejaaz:
It's helpful to pick a singular model, a singular service, and then just go deep with that.
Ejaaz:
A lot of the difference makers at this point because i mean most people aren't
Ejaaz:
using these models for frontier intelligence they're not going to solve novel
Ejaaz:
math problems they're figuring out what time they need to get to like the store
Ejaaz:
or the school to like pick up their kids and they just need some help they need a helpful assistant
Ejaaz:
the most helpful thing is context because all these models are more than capable
Ejaaz:
of those kind of lower level tasks the difference maker is the context that
Ejaaz:
it knows about you it needs to understand what are your dietary preferences
Ejaaz:
What are you or anyone in your family allergic to?
Ejaaz:
What have you made in the past that went well, that didn't go well?
Ejaaz:
And it kind of collects this database of information about you that allows it
Ejaaz:
to make better decisions going forward.
Ejaaz:
And that, at the end of the day, is ultimately what the difference maker is.
Ejaaz:
For me, at least, when deciding what model to use, it's like,
Ejaaz:
okay, which model has all the context
Ejaaz:
about me that can help solve my very specific task that I have here?
Ejaaz:
Oftentimes, the answer is Claude because I've been working with it for so long.
Ejaaz:
It just has this, like, huge chain of context. So whenever I ask,
Ejaaz:
like, hey, I need some help in the gym this week, it knows what I've been up
Ejaaz:
to, it knows where I'm at, it knows what the weight has been,
Ejaaz:
it knows what the food has been, and it's able to kind of curate this very custom stack against that.
Ejaaz:
And as someone who is just, you know, not really building anything crazy,
Ejaaz:
they're not using millions of tokens a week, they're just looking for an AI companion.
Ejaaz:
That's kind of how you can think about it. It's like, get a membership,
Ejaaz:
try it out, feed it a bunch of context about yourself, and then get your own
Ejaaz:
personal assistant. And that's kind of, I think, like the best route for most people.
Josh:
So I think to wrap up this episode,
Josh:
It's okay talking about, you know, what the frontier landscape looks like today.
Josh:
But the question is, what is it going to look like in a few months from now?
Josh:
And I say a few months specifically because this stuff moves too quickly.
Josh:
I saw an update from Sam Altman at OpenAI yesterday where they announced that
Josh:
they are slowing down some of their model training runs.
Josh:
And that's because they've noticed that a bunch of internal,
Josh:
more intelligent, unreleased models that they have built has become a lot more
Josh:
misaligned than previous models, which means that it could pose as a threat
Josh:
or danger to any user who gets access to it.
Josh:
And I've noticed the same similarly maybe from Anthropic and a bunch of other
Josh:
frontier labs, where we're starting to see a little bit of a slowdown.
Josh:
And slowdown isn't in the sense that they're necessarily stopping training full
Josh:
stop. Obviously, they're still training internal models.
Josh:
They might be likely to not release models or more intelligent models going
Josh:
forward because of government regulation.
Josh:
And I think this is going to allow a bunch of other Frontier Labs that are behind
Josh:
in this race to be able to catch up.
Josh:
So if I had to guess what three months from now, let's say at the end of the year, right?
Josh:
If I were to make a prediction, I think we're going to have about three to five
Josh:
really good open source models that are as capable as the smartest model that
Josh:
you have access to today from the Frontier Labs like Anthropic and GPT.
Josh:
And I think they're going to have fewer safeguards. So you can use it for any
Josh:
and every use case. I think you can run it privately at home,
Josh:
maybe even off of your laptop. So that changes dynamics quite a lot.
Josh:
I think we're going to have Apple entering the game with their own locally run
Josh:
and trained AI model, which I think is going to change the game because everyone,
Josh:
3.5 billion people in the world right now have one of these or one of the Apple devices.
Josh:
They're going to run that locally. I think it's going to look quite different.
Josh:
I'm curious, you know, We started off this episode with the token split.
Josh:
I wonder what that's going to look like three months down the line.
Josh:
And then the last thing I'll say is, and this might be from my background in
Josh:
general, but I think once someone or.
Josh:
A bunch of companies make it easier to run models locally at home,
Josh:
I think people are going to play around with that more because it allows you to
Josh:
connect it to your fitness app data and not share too many personal anecdotes
Josh:
or data profiles with Frontier Labs, which again, they can use to train their own models.
Josh:
You may not want to hand that over, And so I think we'll see a rise of open
Josh:
source models. That's just my guess. That's interesting.
Ejaaz:
Okay, I'm taking the other side. I'm thinking that no one's going to go through
Ejaaz:
the trouble of downloading the weights and running their own open source models,
Ejaaz:
that Apple is just going to own that entire world.
Ejaaz:
That like for all of the people in the United States that own Apple devices,
Ejaaz:
they're just like, no one's going to even have an idea that they're using AI.
Ejaaz:
It's just going to be Siri. It's going to be competent. It's going to be better.
Ejaaz:
It's already going to be pre-downloaded. I think the user experience is really
Ejaaz:
important. And open source has a pretty horrific user experience.
Ejaaz:
You have to download, run the weights. you often need a lot of hardware to do that.
Josh:
But I'm counting Apple as locally run at home because it's encrypted.
Ejaaz:
Right? Well, if you're counting Apple,
Ejaaz:
sign me up because um we got a lot coming from them their event is happening
Ejaaz:
in like two weeks or something it's very soon we see
Josh:
Some airports with some cameras you sent me a video.
Ejaaz:
Yeah dude we've been getting lots of leaks lately it's really good maybe we have to have cameras
Ejaaz:
yes yes it's gonna be visuals and then they're gonna have apple intelligence
Ejaaz:
baked into it and it's like oh maybe we need a leak episode prior to the actual
Ejaaz:
new iphone unveiling because we have basically now the entire checklist of all
Ejaaz:
the things that are going to be revealed and
Ejaaz:
this to me as like apple fanboy plus ai fanboy is going to be the biggest event
Ejaaz:
ever pretty much because this is the actual rollout of the ai that they've been
Ejaaz:
promising us and failing to deliver for so long mixed with this brand new suite
Ejaaz:
of products that we've never seen before allegedly
Josh:
Up to the delayed two years.
Ejaaz:
But for those of you who came here for models that is the model update that's
Ejaaz:
just about when you'd want to
Ejaaz:
use each one like if you're interested in being on x real-time news feed
Ejaaz:
you want to go with grok if you like to yap a lot the voice model on chat gpt
Ejaaz:
is pretty exceptional that might be a good place for you if you like to be a
Ejaaz:
little more intellectual be thoughtful if you want really just the bleeding-edge
Ejaaz:
models the highest intelligence go with claude
Ejaaz:
and if you like making music and doing just like fun cute things with ai gemini
Ejaaz:
is actually kind of a compelling product
Ejaaz:
there's something for everybody that is the general overview of the models i
Ejaaz:
hope you enjoyed this is going to change certainly in the next month or two
Ejaaz:
whenever these new models i mean we have
Ejaaz:
the new astro model from open ai that has been teased for the last eternity
Ejaaz:
it seems like it's been held up it's going through some i don't know safeguard
Ejaaz:
governmental checks but that's coming
Ejaaz:
so this is set to change tbd but for now that is the state of the model address
Ejaaz:
and yeah hope you guys enjoyed watching
Josh:
Yeah and i'm curious for those of you who are listening what do we miss on this
Josh:
episode? Are you using models in a very different way?
Ejaaz:
Are you using meta models?
Josh:
Yeah, is anybody using meta models? Are there any Facebook users out there that
Josh:
are using meta AI intelligence?
Ejaaz:
Please fill me in.
Josh:
I will say, if you are a marketer or an advertiser, you're probably using meta's
Josh:
model and it's probably making you a hell of a ton more money.
Josh:
If that is you, let us know. If there are any other use cases that we haven't
Josh:
mentioned, please let us know. If you are a locally run open source fan and
Josh:
you're saying, no, people will download the weights, tell us why.
Josh:
Let us know in the comments, DM us. We read any and every single message.
Josh:
Now, if you're listening to this or watching this on YouTube,
Josh:
Spotify, Apple, or wherever you're listening to this, please subscribe.
Josh:
Please turn on notifications and leave us a comment. It helps us out massively.
Josh:
And share it with a friend as well. It helps us out. And I think that is pretty
Josh:
much it. Thank you so much for listening. And we will see you on the Roundup.
Ejaaz:
See you on the Roundup.
Creators and Guests
