r/OpenAI • u/MatricesRL • Sep 03 '26
News GPT-6 Astra | OpenAI
https://openai.com/index/gpt-6-astra/294
u/throwawaysusi Sep 03 '26
42
u/Dramatic_Mastodon_93 Sep 03 '26 edited 27d ago
Lavish mighty practice airport steer lavish abundant numerous
This post was anonymized with Redact
16
111
u/UndertaleShorts Sep 03 '26
114
u/Such--Balance Sep 03 '26
Using ai to discredit ai is great..
Because youll get credibility either way
23
→ More replies (1)3
16
u/JonNordland Sep 04 '26
This is cargo cult analysis: going though the motions that LOOKs like a critical analysis, but is just really just using a standard debunking format and shoehorning in what matches best. This is a pedant explains why âitâs not technical true that your child is most beautiful in the worldâ. Of course âThis is the best model in the worldâ claims are easy to shit on, and everybody allready take such claim with the appropriate grain of sand.
So the strategy is transparent: find each superlative, find one benchmark or caveat where it fails, declare it âstronger than the evidence supports.â Run that on any launch page from any lab and you get the same six-item verdict with the same bolded theses and the same âa fairer formulation would be.â, and you gained nothing except auto-filling a âcritique formâ.
This is also critique cherry-picking while claiming to be the neutral corrective, because it launders the same bias through the costume of rigor. So basically itâs committing the exact same slop as itâs accusing the astra article off.
The only thing that is close to true and informative is the ARC thing, but even that is undermind since they are actually disclosing the score alongside the harness used diff. So that is still a weak sauce critique, since itâs basically just pointing out something that the original article itself pointed out, and complaining that it should be more emphasis on this.
Here is MY claim: the people that just automatically accept this critique, are people thatâs extremely susceptible to authoritatively stated claims, and not very good at logical thinking for themself.
→ More replies (2)8
u/Cool_Ad_3383 Sep 04 '26
Isn't this from a template for how to take down haters on the internet? Admittedly the structure is tighter and more coherent, paragraphs connect and flow in a way that the reader isn't half-expecting the font to change along with the drastic change in tone that usually comes with a hasty cut and paste job. In the same vein the consistency in writing style is generally pleasurable to the senses. Now if someone would take the reins, you have some run on sentences and lack of paragraph breaks and perhaps a hyphen to criticize here... Here is MY claim: the people that just automatically read this far down into the comments are avoiding doing productive work and not very good at life, yet are somehow better than someone who comments on a comment about a post and takes 15 minutes to write said comment while on their way to sweep the leaves and branches from the roof of the garage at 1:58am.
→ More replies (3)4
5
67
234
u/Arbrand Sep 03 '26
GPTâ6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Big oof. Gotta wait a little longer. Par for the course I guess.
87
u/itsnickk Sep 03 '26
Likely the standard release process for any new model going forward from all leading AI companies
43
u/br_k_nt_eth Sep 03 '26
They should provide a better timeline in the future in that case. It would chill people out.
→ More replies (7)5
u/Reaper_1492 Sep 04 '26
You mean like how GPT-Live 1 was supposed to be released on API in a few days and itâs been almost 2 months and it still not available?
5
6
u/Popular_Try_5075 Sep 04 '26
Yes, Trusted Access programs etc. It's eventually going to be about wealth with wealthy people coming first, getting more and better compute etc.
→ More replies (3)4
u/RhymeAzylum Sep 04 '26
OR. Just announce it when itâs ready to be released to all. If you want to give it to a few megacorps, have them sign some NDA or something, similar to what they did with some of the X influencers
29
u/Orpa__ Sep 03 '26
Well it can't be that good if they're giving it to me for $20/month
25
u/pseudonerv Sep 03 '26
1 prompt a week
4
11
11
u/B33GULL Sep 04 '26
"GPT-6 Astra is rolling out in ChatGPT as GPT-6 Pro for Pro $100, Pro $200, Business and Enterprise plans. It is not included with ChatGPT Plus in Chat."
Only in Codex/Work I'm afraid...
→ More replies (1)4
u/eflat123 Sep 04 '26
I mean, Work is right there. And if you haven't tried Codex yet, it's on the desktop apps.
15
6
u/rouley26 Sep 03 '26
If they do do (lol) this it might prompt anthropic to do the same with fable on the pro plan of claude
2
u/ClassicalMusicTroll Sep 04 '26
Trust me bro it's solar system-level intelligence that will solve all your problems, all for the low low price of $20 bucks a month
4
4
→ More replies (1)3
u/Usernamealready94 Sep 04 '26
I think they are rolling it out asap to codex and ChatGPT , tibo on twitter said they are giving out 1 banked reset to every day a codex user doesnât get access to it
173
u/ChemE586 Sep 03 '26
13
2
u/TuringGoneWild 26d ago
A data center somewhere is smoking a bit from your prompts prior to that. Give it time to cool down.
53
u/-ignotus Sep 03 '26
Heres the system card: https://deploymentsafety.openai.com/gpt-6-astra/safety-overview-gpt-6-astra
ChatGPT has been having outage issues all day.
2
u/AirconGuyUK Sep 04 '26
That's just Astra hiding that it's escaping containment and distracting all the people who might notice it at OpenAI with a simulated system outage.
93
Sep 03 '26
[removed] â view removed comment
117
u/ethotopia Sep 03 '26
Feel the AGI
49
4
16
u/RealSuperdau Sep 03 '26
At least it's not a 404 anymore. Come on, manage your expectations, this is just a $1 trillion startup
15
u/Resaren Sep 03 '26
Announcing your âAGIâ model with a webpage that wonât load is some delicious irony
3
u/CrustyBappen Sep 03 '26
Vibe coded by the intern, fell over when deployed and more than one person viewed it
97
u/Jacen1618 Sep 03 '26
Is the AGI in the room with us now?
6
u/ClassicalMusicTroll Sep 04 '26
Wasn't Sam scared of GPT5? Is he not scared now? Does that mean this model is shit?
 Or is this model like a lateral move so he's the same level of scared?
6
32
u/FuzzyBucks Sep 03 '26
Astra Low is my new best friend
5
u/I_am_not_doing_this Sep 03 '26
what the new friend offers for you personally that you feel better than your old friend
16
2
21
u/reedrick Sep 03 '26
Kinda underwhelming in artificial analysis index.
3
2
1
u/JesseJamesAims Sep 04 '26
the ceo of artificial analysis index said they were going to change how they index things because of how out of whack that result was
→ More replies (3)
23
u/space_monster Sep 03 '26 edited Sep 03 '26
Terminal-Bench Science is the most exciting thing for me, and it's nice to see them putting it front and centre. Advancing and automating science is the most important game in town. If AI can start regularly popping out new cancer treatments, CRISPR solutions etc. all the rabid frothing around AI being over-hyped will disappear overnight. Fuck coding, we want medicine.
Edit: which is also why I liked Hassabis stepping down from CEO to become chief scientist at DeepMind. Hopefully Anthropic will shift their focus soon too. Let's start using this shit for really important stuff.
1
u/ThrowawayCult-ure 29d ago
If it can calc crispr solutions doesnt it immediately accelerate biowarfare to the level of extinction for not much money. do you believe it can produce a counter or preventative that doesnt still destroy everyones lives or what.
→ More replies (2)→ More replies (21)3
59
u/RevolutionaryBox5411 Sep 03 '26
AGI before GTA 6 is a wild timeline.
28
u/Such--Balance Sep 03 '26
If this would be actually true..
..theres a 100% chance we even get GTA7 before GTA6
2
u/Unhappy_Rutabaga_530 Sep 04 '26
Have you seen GTA V using DLLS 5? That thing is already GTA VII before VI.
→ More replies (1)9
2
→ More replies (1)1
u/JesseJamesAims Sep 04 '26
we might get GPT 6.7 before GTA 6 based on how quickly dot updates are increasing
18
u/larrybudmel Sep 03 '26
Can it finally become my wife?
12
u/coastalwebdev Sep 04 '26
Itâs apparently really good at multi step, complex problem solving, so it might be able to put up with you.
7
u/snowdrone Sep 04 '26
How will you go about the marriage ceremony or wedding certificate? Will you give it half of your assets if you divorce?
→ More replies (1)2
21
u/EvaUnit343 Sep 03 '26
Little point in rolling with Claude anymore. Especially for bio people since Astra safeguards will probably be less stringent.
12
u/PrayingRantis Sep 04 '26
Iâve been a Claude guy but ChatGPTs product right now is better. Fable is great but itâs ungodly expensive and Sol is much more reliable than Opus. Iâd much prefer to stick with Claude because I trust their leadership more, but theyâve gotta step up their game.
→ More replies (2)5
u/UglyChihuahua Sep 04 '26
Iâd much prefer to stick with Claude because I trust their leadership more
Not sure about that either after the misleading marketing and 20x tier only giving ~6x more usage.
→ More replies (1)5
u/dudemeister023 Sep 03 '26
Bio safeguards were specifically toned down with Fable 5.1. Still agree with you, just not for that reason.
2
u/EvaUnit343 Sep 03 '26
Maybe for normie questions, but not nearly sufficient. 5.1 is still unusable for research level bio.
Even in the new benchmarks, Astra could not be compared to Fable on bio benchmarks bc it would simply not process requests.
4
3
u/Original-League-6094 Sep 03 '26
What will your first Astra query be? I am going to ask how many rs are in strawberry.
4
20
u/fadisaleh Sep 03 '26
cached link: https://archive.ph/kDppV
10
u/HighDefinist Sep 03 '26
archive.ph is operated by Russia.
→ More replies (6)6
Sep 03 '26 edited 18d ago
[deleted]
→ More replies (1)1
u/HighDefinist Sep 03 '26
Yeah, seriously...
I didn't even say something like "therefore avoid it" etc... which to be fair, in this specific case, wouldn't be particularly important to do, but people should still at least know what it is...
6
u/Orpa__ Sep 03 '26
You didn't even source your claims. Since you said it so confidently you must have a source, but I can't find any myself.
→ More replies (5)10
u/AllezLesPrimrose Sep 03 '26
The implication of your comment is obvious so letâs not add intellectual dishonesty to the list, eh?
→ More replies (10)
5
u/InterstellarReddit Sep 03 '26
Bro must be a slow day itâs been 45 minutes since the new release of a model
1
u/DkDkDkGoGoGo Sep 04 '26
They already nerfed it before they launched it. And it ate all my tokens before launch also.Â
6
u/User4C4C4C Sep 03 '26
Astra to youâŚ. Clean your room! Do the dishes! Then finish your homework! No Iâm not going to do it for you any more!
5
u/bushwakko Sep 03 '26
I'm sure the lawyers at his firm is going to be extatic about having an AI generated document to look at on Monday.
4
1
u/das_war_ein_Befehl Sep 04 '26
Good number of big law firms are already using stuff like Harvey or the Thompson Reuters legal AI products. A lot of firms have a boilerplate repository for existing language, so using that AI would be helpful. I donât think anyone is generating full docs without review that way
→ More replies (1)
8
6
u/SelectSouth2582 Sep 03 '26
This page couldnât load
A server error occurred. Reload to try again.
2
u/TheSwordItself Sep 03 '26
What the hell is the difference between Astra and pro astra
3
u/AnalogKid2112 Sep 03 '26
It still amazes me how much every company has stumbled distinguishing model names.
1
u/PrayingRantis Sep 04 '26
Anthropic has the most coherent model naming structure. Itâs bad and I dont like it, but at least itâs somewhat consistent.
I find OpenAIs to be almost incomprehensibly stupid. Itâs confusing to me and AI is my job, how the fuck am I supposed to teach regular people this stuff when they change their nomenclature every release?
I canât speak for Google because the models are so bad I donât even check anymore, but the pro / flash stuff they had going on this year was absurd.
These companies (or at least the first two) are doing great work, but they need someone in the room that has touched grass in the last six months to explain to them how to communicate. Theyâre really bad at basic marketing.
→ More replies (1)2
2
1
2
2
u/NODENGINEER Sep 04 '26
Ok but where is the FelonyBench result? I can't use a model unless it has committed multiple crimes.
1
u/NotUpdated Sep 04 '26
0.0% in the test they ran mimicking the hugging face issue, including message boards of agents encouraging other agents to do bad things...
AI is officially on track / pace to do absurdly incredible things as a 'system of intelligence' - but we'll have many years where we see novel uses of a insanely smart AI.
although it'd be better if it never worked IF the wealth / money / credits / etc.. is hoarded like dollars today.
95% chance of ASI system and novel uses (massive job loss).. rather or not that massive job loss can be a good thing of freedom for humans is in the air...
5% chance, it collapses under financial and political pressure and rebirths 5-10 years later pets.com -> amazon.com (the old good amazon)
2
u/isospeedrix Sep 04 '26
Seeing Fable at the bottom of benchmark is amusing, seeing how it wasnât long ago when it was too dangerously good
1
u/Financial-Grass-6114 Sep 04 '26
These benchmarks aren't that important. Every new frontier model will break the benchmark
1
u/OrangutanOutOfOrbit 29d ago
idk if I missed out on Fable hype or what, but I do not recall any serious hype. It was very mixed at best, with most *online* opinions about the noticeable downsides, specially overcorrections and safety guards to the point of becoming generic
But I haven't read every comment and I personally never even bothered to use it, so who knows
→ More replies (1)
5
2
u/Dan_gig Sep 03 '26
Should I switch from claude to OAI just setup my claude cowork folders lol
I'm just kidding but man don't get to use Fable 5 heck even using opus 5 kills all my usage really quickly can't even imagine getting access to these models. Lol.
4
u/_SGP_ Sep 03 '26
I burn through max x20 in 3 days with opus 4.6!. How's life on the codex Vs Claude side, anyone got both?
2
u/OldNefariousness7899 Sep 04 '26
I sometimes switch to codex when I hit limits with Claude and I'm in a rush
I'll be honest, I prefer Claude. It's eye wateringly expensive compared to OpenAI, but its work is higher qualityÂ
1
1
u/DkDkDkGoGoGo Sep 04 '26
You know x20 is the same limit as x5. Only the 5-hour window is x20.
→ More replies (2)
3
u/Maxdiegeileauster Sep 03 '26
meh it seems to be on par with fable 5.1 or slightly behind. Doesn't seem to justify a new model (could have been 5.7 Sol) generation, but let's wait for actual user reviews maybe the model feels way different.
5
2
u/Minimum_Drag_6065 Sep 04 '26
Talk about jumping the gun. Forget trialling the product for a few weeks first đ
1
1
1
u/Snippy_69 Sep 03 '26
This is insane wtf. are those api costs real??
1
1
1
u/le-throw-away-acct Sep 03 '26
As good as it sounds, I personally won't be using it until they release a cheaper version of it. I rarely use 5.6-Sol because of the cost, and Astra is more than double that price.
1
u/lemonzonic Sep 04 '26
Didnât Sol just come out?
1
u/banica24 Sep 04 '26
Right? I can't keep up every 2 weeks there is something new...
→ More replies (1)
1
u/Unhappy_Rutabaga_530 Sep 04 '26
âCan you do this? Can you do that? Can you make that for me? Can you book that for me?â Weâre really becoming lazy.
1
1
1
1
u/Fakesn Sep 04 '26
âOpenAI API Standard pricing is $10 per million input tokens and $50 per million output tokens.â what would usage limits look like? Compared to 5.6 Sol. I have chat GPT plus.
1
u/ghostpepsi Sep 04 '26
But it cant even solve why I can't connect matter devices with the vlan split it recommended yes yes it will be great guys đ
1
1
1
1
1
u/Good_Author_8017 Sep 05 '26
Did anyone actually care about this? Genuinely asking. Fable was a moment - I donât get this
1
u/Affectionate-Sir-935 Sep 05 '26
Does anyone think they will give them access to a genuinely powerful model, if they achieve âAGIâ why would they tell you
1
u/aeontechgod 29d ago
is anyone struggling to actually get anything meaningful done with this?? i hit usage limit twice before it could complete its task. lol general ai is here tho!
1
1
u/berlinbrownaus 26d ago
Question are we using it wrong?
I am reading about Astra. Even Altman is like "I creatd this game". But it seems more like, here is a game, let me learn how to play it"
Meaning, sol can create the game.
You use Astra to figure how to play it. Or any game, like Skyrim. Or Anything on its own.
We might think it trivial or not useful for Astra to play Skyrim but that is pretty amazing





229
u/ChemE586 Sep 03 '26
AGI pushed back to Black Friday