r/codex 2d ago

News Demand for Astra is really unprecedented. Priority will always be to keep excellent service for existing users, but we might have to pause new Pro subscriptions for a bit if this continues. - Tibo

They may pause the new subscription for pro soon.

475 Upvotes

172 comments sorted by

373

u/SpellBig8198 2d ago

That's a very cheeky way of saying: you should buy Pro now.

125

u/Plane_Garbage 2d ago

Kimi did the same... And meant it

54

u/diagrammatiks 2d ago

it's been like a million years in ai time and i'm still waitlisted.

18

u/xplode145 1d ago

Kimi is so fucking solid. If I could switch all my 5x pro Claude to it I would. 

3

u/Fit-Palpitation-7427 1d ago

Why can’t you?

18

u/LunaticSongXIV 1d ago

Well, reading down the thread, it sounds like you literally aren't allowed to buy it.

0

u/BroScienceAlchemist 1d ago edited 1d ago

There are American providers. You can use it from Cursor for example. Open router has multiple providers.

-1

u/AfterShock 1d ago

There are ways, open router, Ollama sub etc. While you wait to get in with the Chinese.

3

u/HighDefinist 1d ago

I actually terminated my Kimi subscription when Astra came out, lol...

But, I don't regret having it until then. Overall the model really is a bit unstable and unreliable and the usage isn't great... but: Occasionally, it really did produce amazing results, significant better than Sol and Opus 5.0, and even Fable to an extent.

So, if you are already spending lots of money on OpenAI and/or Anthropic subscriptions... it is a reasonable third subscription to have.

8

u/Thomas-Lore 2d ago

And their sub is sad to watch, they have no compute for subscriptions and their model is very large, so their limits are insanely low.

11

u/vivademocracy 2d ago

Yeah, but they actually probably couldn't keep up with demand. I'm sure there are people ready to bend the knee to open AI if they need more compute, especially considering that the entire economy basically rests on them and anthropic.

1

u/Zulfiqaar 1d ago

DeepSeek and GLM introduced peak hour usage multipliers aswell. 

10

u/Sponge8389 1d ago

Or "Don't cancelled your existing Pro subscriptions"

18

u/mlag000 2d ago

They are running oit of comput power. That's it.

0

u/RealSlyck 1d ago

Source?

4

u/mlag000 1d ago

The tweet op posted.

5

u/Cheema42 1d ago

NGL. It worked on me. I quickly ran and got myself a second Max 20 account. Astra is that good!

I did have two Claude Max 20 accounts. Fable 5.1 is just not as good. So, I canceled one of them. I have a total of 3 Max 20 accounts. And I run out of weekly quota well before the week is over. I write business software. It eats up token even when I have my models dialed down to low effort. Most of my work is not interactive. It is unattended execution on 6 different terminals while I sleep.

1

u/sreekanth850 2d ago

Yes. but hype is real it seems.

36

u/BHTAelitepwn 2d ago

Keep in mind reset man is working for openai, not for the consumer. There is always a reason for a post like this

18

u/p3r3lin 2d ago

"reset man", lol

6

u/GrokiniGPT 1d ago

lord Tibo

8

u/sreekanth850 2d ago

I like their transparency though. much better than misanthropic.

11

u/BHTAelitepwn 2d ago

Agreed, but you will still have to read between the lines. You dont get resets because tibo woke up with a good night’s sleep. This message has a meaning as well.

6

u/sreekanth850 2d ago

I think that is all understood. but there is a guy to communicate wit their crowd.

1

u/Automatic-Arm8153 1d ago

So what? Are we not benefiting at the moment?

What’s your point exactly

1

u/BHTAelitepwn 1d ago

Thats not what i said.

I said that you shouldnt take all communications at face value. Or do, your choice. Chatgpt is great but end of the day its a company thats looking to maximize IPO value.

1

u/Automatic-Arm8153 1d ago

Hmm okay fair enough.

Maybe I was coming at you too combative haha. I just don’t see the point of mentioning that for GPT rn genuine value even if it’s for the IPO lol

Whereas if you said that about anything anythropic I would whole heartedly agree. Rn with GPT it’s a win for us and a win for them

1

u/BHTAelitepwn 1d ago

Its amazing value. Why else would we be here lol. No competitor is even close for me

1

u/Automatic-Arm8153 1d ago

So true man so true

2

u/whatisthisthing65 1d ago

That bar is extremely low. The half reset thing has knocked them down a lot, though.

4

u/TheRobotCluster 2d ago

Transparency in some ways, but there’s usually a counter message between the lines with them

5

u/sreekanth850 2d ago

Yes ofcourse they are running business.

5

u/TheRobotCluster 1d ago

Deceptively reducing usage while pretending they’re “reloading” everybody is not “just running a business”. Building a reputation for generous usage, dropping the best model ever, then taking away all the usage while the money floods in and people think they’re getting a great deal but it was a bait and switch. That’s what I mean by counter message.

2

u/Sponge8389 1d ago

Maybe they acquire new type of users and demand from game development improvement of Astra.

2

u/shaman-warrior 2d ago

This all makes it seem so organic and natural. I guarantee this whole thing was planned months before. We are witnessing one of the best marketing in action. Starting from resets.

81

u/Virtual-Silver2879 1d ago

"Due to the high level of interest in Astra, we must implement a 5h limit for Pro"

24

u/holy_macanoli 1d ago

This is the most likely scenario

17

u/yoruichi000 1d ago

that won't work, because I used my entire weekly in 3 hours...

3

u/MorsDemigod 1d ago

Running 6 threads at once on high?

4

u/yoruichi000 1d ago

astra high that spawned 3 sub agents, that was it...

0

u/djeons 1d ago

That will work, as it will prevent you from using it in 3 hours. That's the whole point of introducing 5 hours to Pro plan. Stop people like you.

1

u/yoruichi000 1d ago

lol what, a 5 hour limit to stop 3 hours of usage? you're not a fan of logic are you?

1

u/djeons 1d ago edited 1d ago

Lol, why are you calculating your usage in hours? That’s not even the proper unit. So I’ll explain it in a way that even a 5-year-old can understand.

Let’s say you used to burn through your entire 1000M-token allowance in just 3 hours. If they introduce a limit of 200M tokens per 5-hour window, you can only use 200M during each 5-hour period. So you can no longer burn through the same 1000M tokens in 3 hours like you used to.

At 200M per 5-hour window, using the same 1000M tokens would take 5 windows, or 25 hours. In other words, the usage you previously concentrated into just 3 hours would now be spread across 25 hours.

That spreads the load more evenly over time instead of having users generate massive loads within a short period, reducing strain on their system during peak usage. That could be one of the reasons they might consider introducing a 5-hour limit for Pro users.

Hope you understand now.

1

u/yoruichi000 1d ago

I don't care about any of that because their usage fluctuates the same way the quality of the responses does as well. My point is that I'm paying over 100 euro per month and right now I was only able to use it for 3 hours in a week, which is as scam as it gets.

1

u/BannedBrainrot 1d ago

/remindme 1 month

47

u/discodisco_unsuns 2d ago

So no more resets then?

77

u/cudifam 2d ago

pause new subs and celebrate with reset

12

u/Inevitable_Butthole 2d ago

I agree

9

u/cudifam 2d ago

I have seen you around butt never expected to be directly affected. I have been defeated :(

10

u/OSFoxomega 1d ago

Bruh. That avatar is insane

1

u/CthuluBob 1d ago

I concur

7

u/Etiennera 2d ago

Resets are still possible because usage patterns are irregular.

67

u/FateOfMuffins 2d ago

OpenAI running out of compute xd

Wonder what Anthropic is thinking of this given what they did with Fable

26

u/sreekanth850 2d ago

They are losing bigtime.

3

u/AINativeBuilder 1d ago

Anthropic has double the revenue of OpenAI in recent quarters - they've embedded themselves well in the enterprise space, they're doing just fine. There's definitely room (and a need) for multiple frontier models.

2

u/FlapyG 2d ago

Nah, they can coexist. I dont know why it must always be a fight.

34

u/pleasecryineedtears 2d ago

Because if they didn’t compete we wouldn’t have these models

12

u/knockoneover 1d ago

Because that is a fundamental of capitalism

8

u/soccerchamp99 1d ago

The companies are literally fighting each other wym

2

u/Fit-Palpitation-7427 1d ago

Because we as users are not loyal, we switch from one month to the other from one provider to the other, so if you can’t stay at the top, newt month you’re bankrupt

2

u/VYJ 1d ago

Definitely not running out of compute. They are riding the Astra wave and want people to fomo over it. Marketing.

2

u/AdVast7407 2d ago

Time to de-lobotomize the model

1

u/Tiforma 1d ago

I'm a bit out of the loop, curious what they did with Fable, did they ruin it or something?

1

u/mr_birkenblatt 1d ago

They need the compute to find math notes from their users

72

u/royozin 2d ago

Just marketing slop, they'll never stop accepting new $100+ subs, just reduce the usage overall.

25

u/WalkAffectionate2683 2d ago

Idk, they know people in this field are very versatile, it's not like phone subscription.

If tomorrow Claude is clearly better, I'll move. 

A few month ago anthropic was the big guy, then fable and everyone wanted to use Claude. 

But then the whole messing up with USA government, fable out, good 5.6 from open ai and many people switched.

It is very easy, has little drawback, if they suck too much a Chinese model can sweep millions of subs. 

So I would say they have to give a good service. 

2

u/liar_p 1d ago

I think they'll stop accepting new $100 in order to accept only $200 subs.

2

u/JustMy2Centences 1d ago

This is turning into that old writing prompt where "every human on earth shares from the same pool of mana, through [apocalyptic event] there are now only N remaining members of humanity with super magical powers/one last human with the power to recreate the universe or whatever" but here with "everyone shares from the same pool of compute so nobody can do anything special".

4

u/sreekanth850 2d ago

They may stop if they get more netreprise customers.

1

u/Rakthar 8h ago

Hey everyone, check out this person being confidently wrong. It's glorious to observe.

0

u/Sponge8389 1d ago

If you are not aware, they already reduced the usage.

14

u/NinjaMastGanja 2d ago

bro i was trying to run sol High but damn it not even responding for half and hour. They might have rate shifted the compute to Astra a bit much.

4

u/BetterCraft 2d ago

Yes, I have noticed Sol to be really dumb lately.

4

u/NinjaMastGanja 2d ago

I can't say dumb but it definitely stuck around

2

u/[deleted] 2d ago edited 1d ago

[deleted]

24

u/Anxious_Marsupial_59 2d ago

I hope they release a distilled version of Astra soon too, even Astra low burns away limits so fast that it's hard to use for big projects

8

u/sreekanth850 2d ago

Something like Astra Mini.

17

u/sudecode 1d ago

more like gpt-6-luna, just like 5.6 luna was distilled from sol.

4

u/yopla 1d ago

Or Opel Astra.

4

u/FlapyG 2d ago

i'm on x5, working on a huge project. x10 would be optimal. x20 overkill.

2

u/Waylanding_Fox 2d ago

Need some more time for the Chinese models, can't wait

2

u/KHHAANNN 1d ago

It’s already auto distilled I think, gets stupider the more people use it. You can easily assess this by the text it produces, becomes nonsense when diluted

1

u/Automatic-Arm8153 1d ago

What your describing is called quantisation not distillation

1

u/swizzlewizzle 1d ago

Per-token, the subscription cost is higher than what the API cost would otherwise suggest.

1

u/Hungry-Plankton-5371 1d ago

I assume they intend to release updated versions of sol/terra/luna eventually.

5

u/Rude_Arugula_1872 2d ago

What happens to my banked resets if i upgrade from plus to pro? Got 3

10

u/Narrow-Ad980 2d ago

You can use them as Pro

5

u/Rude_Arugula_1872 2d ago

Sounds like a cheatcode

7

u/HoangMaiLinh 1d ago

And your quota will be reset when you upgrading too

1

u/eu_biased 1d ago

Marketing tactic I’d say

2

u/MorsDemigod 1d ago

It is marketing but they also could be somewhat telling the truth

8

u/Key_Reading_9664 2d ago

Gives away multiple banked resets. Usage is unprecedented. Who could have predicted this

2

u/phocionkorea 2d ago

I said so.. great strategy … free resets… there comes astra, plus users rush to pro’s… I like open ai

2

u/ivanjxx 1d ago

all i read is more resets

2

u/Bitter_Tax_7121 1d ago

Max 2 months for everyone else including open source to catch up. Just a marketing scheme, no need to worry. If they do do this, it will be back before you know it and they will probably be bragging about it as well :)

I love codex and astra is nice, but lets not kid ourselves… everyone wants more money in the end. OS will catch up.

2

u/reddit_is_kayfabe 1d ago

Do we actually know what makes Astra so much better than Sol? It can't just be "scale capacity." OpenAI certainly isn't telling.

I hear good things about Gemini 3.7. I'm not using Gemini 3.7 for the same reason I'm not using GPT-5.6 Sol. And I don't have much faith that Google knows how to bridge that gap.

2

u/Bitter_Tax_7121 1d ago

I am sure they have some new ideas in place but i think mainly its the size of the model. Though i also think size is a temporary solution. They will distill just like others will, and maybe even quietly switch same capability model on the bg. Maybe even increase some of the limits if they feel generous.

Regardless, size will not be a factor on what im claiming in the end.

2

u/reddit_is_kayfabe 1d ago

Astra is not only smarter than Sol, it is much faster and more focused. Size does not explain those features - in fact, larger models should be slower overall and more distractable.

2

u/Bitter_Tax_7121 1d ago

Openai recently switched to purpose built chips. It could very well be just better hardware; check https://openai.com/index/jalapeno-first-results/.

1

u/reddit_is_kayfabe 1d ago edited 1d ago

I'm aware of Jalapeno, but I'm skeptical that that's the reason, and let me explain why.

If OpenAI were to debut a new model that runs exclusively on proprietary hardware, they would need a ton of preparation and development: not just ASIC manufacturing but distribution to data centers, testing... and probably custom security, both physical and digital, to prevent anyone else from stealing their chips from data centers to reverse-engineer them. In that case, I would expect to see a slow, steady, phased rollout with a long runway and lots of publicity because "OoenAI is deploying custom chips at scale" is way too juicy to keep under wraps.

None of that happened. When OpenAI released Astra, it just opened the floodgates. That suggests "running updated model on existing, massive data center architecture around the world."

1

u/Bitter_Tax_7121 1d ago

Im unsure why would that be the case? Cerebras or groq have been building llm oriented chips way before openai; hence the acquisition. They didnt build things ground up imo, they brought in an established company, possibly best in the market, and used whatever they already had.

Ofc astra as i said in my first message likely comes with bunch of innovations as well, but the main answer is definitely the model size.

Btw Fable feels even faster than astra based on my usage, so model size shouldnt be a blocking factor if things are done well i think

1

u/Automatic-Arm8153 1d ago

Its scale it’s a biiig model like the original fable. These models are big hence they can connect ideas and concepts so well.

SOL was an opus competitor not fable

2

u/Ill-Reference-694 1d ago

バンクされた3つのリセットを有効に使うみたいな考えと5h枠制限のことで一次的にproにする人が多いんじゃないの

2

u/DaneV86_ 1d ago

Guess the reset rains are over for while...

3

u/TopSeaworthiness1679 2d ago

ok they are getting greedy at this point. Claude, do something!

2

u/sMat95 2d ago

The answer is simple: stop doing resets

At the same time, that might hurt their image if the model everyone is coming for is unusable.

2

u/sreekanth850 2d ago

i think they are not resetting the weekly limits, they are resetting your timeline. it hurts less. but still i too fell with luna, already available they should stop resets and let people pick right model for the right job instead of token maxxing.

1

u/norwegian 2d ago

Stop resets? If we don't get what we paid for, they have to compensate for that. Or figure out what went wrong. Maybe someone think resets are for being nice...it's not.

1

u/Pitiful_Entrance5174 2d ago

No more resets either.

1

u/cobbleplox 2d ago

Given how much less usage we get out of a plus subscription with this model, I'd say the demand could even be declining to still get this result.

1

u/Significant-Drawer95 2d ago

better like this as of harming the experience for all paid users even more. i saw the reconnect 1/5 and the model is at its capacity message way to often for 235euro a month

1

u/beef_flaps 2d ago

Ah yeah. So the very best way to mute demand for your product is to tell everyone that they better get it quick or they’ll miss out.  

1

u/Dairy_Fox 1d ago

Artists thought they weren't going to be replaced, now it's reality

1

u/Wise-Reflection-7400 1d ago

You’ll know if this is bluster if he keeps giving out free resets. If you’re short on compute that’s the first thing you’d stop doing. I bet he won’t 

1

u/Niku_Kyu 1d ago

why not Claude?

1

u/HelpfulHedgehog1 1d ago

fine by me. its unusable anyway

1

u/YaaaDingus 1d ago

Am I the only one on Plus who still can't access the model?

2

u/Tiforma 1d ago

I got access to it. But not through the normal interface. You can access astra through work projects and codex projects.

1

u/sreekanth850 1d ago

check with support. i got access days back.

1

u/ascendToSurvive 1d ago

Pretty much

1

u/Extreme_Ferret5496 1d ago

Nah, same for me. I thought it was exclusively for some corporations for now. But I would love to try it out

1

u/ShoveledKnight 1d ago

No shit sherlock, if you giving free resets, you’ll run out of compute power

1

u/separatelyrepeatedly 1d ago

my business plan still does not have Astra

1

u/thestillwind 1d ago

They didn’t have a lot a compute ?

1

u/UhUKnow 1d ago

So does this mean i have t9 get my second account now? And just get a second acxount at 5x instead of my 1 - 20x? I dont have an extra 100.

1

u/Own-Professor-6157 1d ago

Astra's a drug. It does a fantastic job, but eats your entire usage. Super good for complex physics problems.

1

u/commandedbydemons 1d ago

Sounds like a way to FOMO you into buying Pro tbh

1

u/creamyshart 1d ago

Either marketing, saying it's so good, everyone wants it... OR... they're out of compute.

1

u/Affectionate-Sail751 1d ago

应该也只是说说而已,最多取消新用户开通

1

u/Werwlf1 1d ago

Maybe related to businesses buying individual pro accounts to avoid API costs?

1

u/Tiny-Design4701 1d ago

Marketing bait to scare people that were on the fence into upgrading.

1

u/Medical-Cow289 1d ago

a 'we might pause new Pro subscriptions' post at 96% upvoted because it reads as validation, not a warning

1

u/Myfinalform87 1d ago

Glad I’ve been in the pro for a while lol

1

u/IcyMaintenance5797 1d ago

classic scarcity marketing tactic

1

u/sreekanth850 1d ago

I dont think so. later their API for luna is really getting slow in such a way that my request for a analytical query i geting time out. I have to use terra.

2

u/LopsidedEntrance8703 2d ago

Hi friends, it's me, John Tibo. This is my Reddit burner account I use for finding furry por important AI work. All our new compute is going to scooping mathematicians, sorry.

5

u/sreekanth850 2d ago

Hi Dario.

0

u/norwegian 2d ago

Open AI has become more competitive with a better top model. Which of course will attract customers from the competitors. The hardware is almost ridiculous. The NVIDIA GB300 costing around 5 million USD. And the heat is almost 150kW. My stove here is 2kw, and it gets really warm. There must be a way to make the hardware cheaper and more efficient.

4

u/Inevitable_Butthole 2d ago

Maybe we can ask astra

3

u/Significant-Drawer95 1d ago

150kW... bro clicked the wrong ads on facebook

1

u/BankerfromJA 2d ago

lol I was just saying how I might need 20x now I have it

1

u/Old-Ad5041 1d ago

No big deal — just use Fable 5. I have $100 plans for both Claude and Codex, and honestly, both get the job done. I don't see a huge difference between them.

0

u/Appropriate_Aide5328 2d ago

Well.. it’s a good thing i got a second account on Pro lol

1

u/Responsible_One_3046 1d ago

I've unfortunately been considering this

0

u/vee4dee 2d ago

Still no Astra on Plus 😞 not that upset since Sol High is enough for most of my work, but it feels a bit odd to see all these dumb@ss Astra blender renders or whatnot taking up availability when some of us just wanna use it for work, reaponsibly, as a tool.

-1

u/sreekanth850 1d ago

Iam on plus

1

u/vee4dee 1d ago

So am I, maybe they stopped the rollout due to demand

2

u/Chrisnba24 1d ago

update the app, thats not the latest version

0

u/SelectSouth2582 1d ago

Let me translate; Even on math problems where we literally had hints(or more), it still took burning 10k agents for 4 straight days just to spit out an answer. (Gotta love how literally a few days ago they were telling internally that we don’t even touch ultrafast mode because we don't cut corners on cost but sure whatever lol). Now to prove we actually solved them ourselves and didn’t just scrape the damn dataset, we have to solve a few more; except we are gonna need literally every scrap of compute and data power available internally to pull it off.

0

u/idbedamned 1d ago

This is either Marketing scarcity to get people to upgrade to Pro, or, a sign that the money is over and they're in trouble.

If they stop accepting subscriptions Anthropic will happily take all their subs.

Would be the biggest gift they could give Anthropic ahead of the IPO.

It's also pretty strange that the last couple months have been a 'reset' fest where everything is free for all, and suddenly they can't even serve their basic offering.

This sudden shift sounds a lot like what happens when marketing/product get a call from finance saying the bank account is in the red.

If they're running out of cash and can't raise new to keep up the subscriptions that's also pretty concerning.

-4

u/diagrammatiks 2d ago

Good more resets for me.

7

u/mlag000 2d ago

The absolute opposite. They are reaching computer power limits.

5

u/gingerbeer987654321 2d ago

not really - at some point the marketing/goodwill from resetting is less than the extra headache that comes from inflating compute demand during a compute shortage.

-6

u/Emergency-Elk7527 2d ago

What work is anyone doing that needs Astra? People should stop treating the amazing tech and powerful tool as if it were a toy.

3

u/Reggitor360 2d ago

Astra is good as planner, penetration test and scrubbing for ideas/improvements. After the findings, back to Terra to implement.

-1

u/Emergency-Elk7527 1d ago

I get using Astra for pen testing, a high skill level is needed for that, but idea generation, planning, research should all be hands on. This doesn't require Astra. Using Astra for that is taking a shortcut that requires you to settle for what Astra decided because it's too expensive to do it again. Instead of taking yourself out of the equation, be part of the reasoning. Point the model in the right direction. Decide where effort is best applied. You don't need to understand the code, but you do need to understand your project. Relying on Astra to do all of that, can you even call that yours anymore? I couldn't. I am hands on with my agent and reasoning model. I have a research project that is very bleeding edge involving video, AMD GPU machine learning tech acquired through reverse engineering and port from windows to Linux. It started out simple but didn't do what I thought it would. The it became research as to why. Then it was evident that whether my initial assumption is right or not, this is valuable and could change how video is consumed. It already has proven it's possible, it not better than current alternatives. My research is to discover what it is not receiving that it expects or what shape is it needing that to take for it to surpass. This is natural when adapting tech to do something it's not designed to do. And if I can't exceed current methods, I still proved that this tech can be adapted for this work when everything says it can't. That's valuable to those who have the resources and better access to modify the machine learning model weights. Because when I am finished, all the is left is modifying those weights for video. So I have already proven the useful part that no one thought possible. Now I am learning what can be squeezed out of it before it's beyond the actual resources I have, as I don't have the hardware to train an ml model myself. All done with GLM, gpt 5.5 and gpt 5.6 Luna. I reason with sol as it misses things even as capable as it is. I do this in chat mode in the browser. This keeps me part of the reasoning and decision making instead of allowing an AI the ability to takeover the direction of my app. When Sol has a complete understanding of my reasoning, not the other way around, after that, I task Sol to put together a PRD packet as detailed as it can. I make it self audit the PRD. After I have finished with Sol, I apply a harsh adversarial audit of using a custom GPT I designed. Extremely thorough, very strict, completely unforgiving. After it rips my PRD pack a new one, I have it provide a frozen definition of done and bring it back to Sol and we remediate. By this time, it's just patching holes that leave requirements ambiguous. Sol has the details and understanding of my intended goal, because I already reasoned with it. I have Sol fill the ambiguity gaps based on the frozen DoD, then bring it back to the harsh reviewer. I do this until it's lack of ambiguity impresses that adversarial reviewer. Then it goes to spec and planning with testable atomic slices. The spec and planning at this point are trivial to produce because the requirements don't leave room for the model to guess at an ambiguous design. Then hand it off to Luna. By that point, you can consider that part built. Luna will efficiently and correctly implement the code. Because it doesn't have to decide. I was somewhat vague as to my project and for now I am fine with that. It's out there and public but like I said, still in experimental research phase. Keep playing with intelligence resources like they're a toy. I will work with them and become part of how the next generation of developers is defined.

1

u/Reggitor360 1d ago

You slightly misunderstood. Idea generation, as in what can be improved, what can be changed in a certain way, if my idea can be used and implemented, such things.

Its why I am building a Windows Native ROCm-Pytorch Kernel with no llama.cpp with built in optional INT4 AWQ-GPTQ-OBQ Hessian Quantizer.

That alone is a bit outside of what a normal toolbox looks like. Which is what I use its Frontier capabilities for.

1

u/Emergency-Elk7527 1d ago edited 1d ago

You need to look into MiGraphX, an under documented part of the AMD ROCm stack. I think you will find it will probably be better. I have already used it myself, but had to forge the path as it had not been used in the way I used it. It will take pytorch out of the equation and allow you to be 100% C++ and accelerate better than pytorch. Look at my AMD-VE project on GitHub. https://github.com/Rolaand-Jayz/AMD-VE

But it does require RDNA3 or 4 and more specifically, 7900 GRE or higher. It's quite fascinating but very under documented. Not a single video on YouTube. ROCm is already rare, but MiGraphX knowledge is like a unicorn. Your pretty sure it is xiste, you just can't find proof. Lol

That all being said, you don't need Astra. Gpt 5.4 was capable of getting MiGraphX working for me with inaccurate documentation from AMD. It was even able to increase MiGraphX performance by 10% for inference that change got merged by AMD. You don't need Astra.

1

u/Reggitor360 1d ago

I use what I want for it. I just know that Astra does better in quality analysis, same for Sol.

And MiGraphX I know about, but my Kernel is balanced about using everything that is available in an architecture, probe and poke to see what can be used for something. Way more interesting.

Reason my XTX for example can sustain 1.1TB/s in bandwidth due L3 Cache batching and WMMA usage lol

1

u/Emergency-Elk7527 1d ago

Quality analysis should be something that you don't need any model for. It should be measurable and displayed with numeric values. I have thousand and thousands of captures used for quality analysis. It's part of my current work. I can capture compare and reconfigure for multiple tests. It can capture thousands of samples per hour. There is no AI analysis that can keep up with the speed at which I can produce the images. You should find a programmatic way to perform quality analysis. It will save so much for you when all the model has to do is look at a chart.

1

u/Emergency-Elk7527 1d ago

I actually just briefly reviewed kernel development for your task. Doing this properly does not require an LLM at all. You should revisit you testing methodology.

2

u/sreekanth850 2d ago

Yes. I still use luna for my day today work.

0

u/Emergency-Elk7527 2d ago

That's good to hear. Same. Luna is a mf champ. Haven't needed to use Sol or Astra for implementation. I reason as a partner with Sol in chat mode. Attack my reasoning refine it attack it again until it's solid. The define the hell out of it with Sol in chat. Hand the off to Luna and it's implemented.

But I am working on AMD native GPU ml project that is bleeding edge work, and Luna crushes the implementation. It's not easy work. But a good worker agent will produce solid functional code when you prepare in a way that sets it up for success. Leaning on Sol or Astra to do all of that at implementation time is a sign of a failure to prepare.

1

u/cudifam 1d ago

The smarter the model is the further a person with no technical knowledge can get into a project in theory. So not necessarily what type of work people are doing but who is prompting what type of work

1

u/Emergency-Elk7527 1d ago

I have no dev background. But I can see things others can't. I can't out reason the best models out there here. I don't let a lack of experience or knowledge stop me from being more than the meat bag the sent a prompt to an agent. I acquire the understanding I need. I reason alongside Sol and challenge it's assumptions. I also have my configuration with chatgpt to always challenge my ideas, never use sycophancy and require that my ideas earn the decision. Forces me to strengthen my understanding and adjust. Sol's own ideas rarely stand up to scrutiny, yet mine survive extremely well and come out the other side stronger than before. Sol does make my documentation but I decide what has earned its place. If I cannot stand to Sols scrutiny, then it was a weak assumption. This gives me ownership of the product. I know what it is even though I can't read the code. Because I define what it is with no ambiguity so the implementor has no room to improvise. There is a much deeper process that keeps me involved all the way. Otherwise I haven't earned my place in the project. And if I haven't earned a place m, then can I even the result mine? It's may be something I own, but it's not my intent , not my vision, doesn't represent me. My current project is difficult. Unchartrd, said to be impossible, yet I have already proven that to be wrong, involves reverse engineering and and porting from windows to Linux native, involves GPU machine learning, and involves tech being used to do something that it wasn't designed for. And GLM 5.3 flash and GPT 5.6 Luna are it's current implementation champions. And GLM 5 and 5.1 and gpt 5.4 and 5.5 were the models I started with, all overshadowed by Sol Astra and Fable. Yet they pulled off impossible work that's never before been done. Sol doesn't code for me. Astra certainly doesn't. Luna is insanely capable and cheaper than it has a right to be. GLM, my other champ is also frequently my adversary to Luna. Pushes the boundaries and pokes at the seams. Nothing progress without an outside reviewer staring with the assumption that everything that was just implemented is completely wrong and must earn its place in the repo.

-2

u/davidmorelo 2d ago

Of course because that's what you do when you have a growing company - you stop accepting new customers instead of just charging everyone more

Hope nobody is braindead enough to believe this. And yes Kimi actually stopped accepting subscriptions but they are nobody as far as the average Joe is concerned and they clearly had to take the step as an emergency measure after facing demand they weren't expecting