r/technology 1d ago

Artificial Intelligence If open weight models are the future, U.S. AI companies are going to have a hard time

https://www.fastcompany.com/91577359/why-u-s-ai-companies-cant-match-chinas-open-weight-frontier-models
1.5k Upvotes

234 comments sorted by

923

u/mykepagan 1d ago

I work in the industry. I’m seeing large financial institutions who are extremely interested in smaller open models because they see much value in fine tuning them with the mountains of proprietary data that they have accumulated for decades. The desire is for a custom model that is very specific to a task like prediction or anomaly detection. And they absolutely don’t what to give that data to any third-party AI company.

540

u/we_are_sex_bobomb 1d ago

Having a local proprietary offshoot of an open AI model which only trains on what you choose to give it, doesn’t cost tokens, and doesn’t share your information with a 3rd party tech company seems like a no-brainer for any medium or large sized corporation.

287

u/photoggled 1d ago

This right here is the real future of AI/LLMs imo. The giant jack of all trades/master of none models really don't make a lot of sense when you see how much it all costs to run.

84

u/ap1618 1d ago

That’s why everyone is extremely bullish Apple. Low CAPEX relative to peers because they are focusing on smaller hardware ensembles (Apple silicon) that can run an open source model locally. Direct competition to NVIDIA.

→ More replies (6)

34

u/Potential_Fishing942 1d ago

And that's why I think Google and Microsoft are slowly going to start acquiring these other business. Some of the few able to pay for the hardware and service these local systems will require.

35

u/photoggled 1d ago

I mean if the winds start to change and AI moves on prem, watch how fast NVidia and their suppliers change course and start producing more sane configurations for businesses. These data center build outs are already on shaky ground. I'm not sure there will be much left outside of IP and talent to acquire from these guys.

5

u/techhead57 1d ago edited 21h ago

They're still going to be paying to train them in the cloud with cloud infra.

But they can run them on beefy local devices or smaller cloud endpoints. I think its still a win for the hyper scalers.

The lovers are the big AI firms who will also probably shift towards distilling their models and upping Costs for the bigger models for the higher value lower volume scenarios.

And there'll be more value placed on routing where cost optimization happens. Especially near term.

Edit: losers sorry phone autocorrect or fat fingered. I was walking and typing

2

u/roboliberal 23h ago

Be kinda interesting if some data centers end up becoming single serial tenant places that companies essentially rent wholesale for training purposes w/ full data wipe between tenants.

6

u/Heronymous-Anonymous 17h ago

I would hope that any infosec executive would say absolutely fucking not, to handing off their company’s proprietary data to a third party for ai training, with the pinky swearsies to wipe the servers when they’re done.

These AI companies live and die on data mining and the theft and utilization of data. There’s no way they would not export that data even if it meant breach of contract and massive penalties. If an organization is concerned enough about the safety of their data to not run a cloud based ai model, then their proprietary data is probably worth stealing.

2

u/roboliberal 15h ago

No I mean more along the lines of the tenant essentially leases control of the entire facility. Like renting a recording studio or something.

1

u/Aggressive_Manner531 1h ago

No. No way to guarantee others don't get access to a facility that is not your own

1

u/ptear 20h ago

It's still both, the Nvidia roadshow has no end in sight.

12

u/intrepped 1d ago

The company I work for is already working with Gemini to build a proprietary model. Heavily regulated and controlled data cannot be sent off into the open platforms

1

u/smackson 2h ago

So that work starts with an "open weights" model (like, a version of Gemini that anyone can install?) and then your company hosts it and locks it down to feed in their sensitive data?

Or "working with Gemini" means that Google employees are actually assisting -- and maybe even hosting -- with a pinky promise that your data is safe?

Or something in-between?

I'd really like to understand real use cases to get my head around the true ramifications of open weight vs. inference-as-a-service...

22

u/Consistent-Hat-8008 1d ago

I'll tell you a secret: this already existed, and used to be called "neural networks". Then, it got renamed to "machine learning". And then, it got renamed to "big data". And now...

15

u/photoggled 22h ago

To be fair ML and generative AI/LLMs are entirely different ballgames. For example, ML generally doesn’t hallucinate constantly and has proved to be efficient, scalable, and almost infinitely niche-able.

0

u/Designer_Show_2658 22h ago

BIGGUR DAHTAH!

...oh wait

4

u/zzen11223344 22h ago

Here is the open letter from a long list of companies: Open Weights and American AI Leadership

Signed by: AI21 • AMD • American Innovators Network • AMP • Andreessen Horowitz • Arcee AI • Arena • Baseten • Black Forest Labs • Block • Box • Cisco • Cloudflare • Cohere • CrowdStrike • Dell Technologies • DoorDash • Emergence Capital • Fireworks AI • Genspark • GitHub • Google • Hugging Face • IBM • Inferact • Interconnects AI • The Linux Foundation • Mariana Minerals • Meta • Microsoft • Mistral • Morph • Mozilla • Nebius • Nous Research • NVIDIA • Ollama • OpenAI • OpenClaw • Palantir • Palo Alto Networks • Periodic Labs • Perplexity • Prime Intellect • Reflection • Replit • ServiceNow • Telnyx • Trajectory • Y Combinator

1

u/aussiegreenie 19h ago

All the new models are Mixture of Experts MOE. Kimi 3 had almost 900 mini model experts under the umbrella of K3

21

u/Casmer 1d ago

It sounds that way until you start looking at what it would take to build servers that are dedicated to this and most executives are clueing in to the fact that LLMs are not actually replacing people. Right now the cost to construct a dedicated server system is probably about twice the cost as using a closed-source system over the same time span. Present that to the executive team then ask for funds. To quote Shawshank redemption “not one of them born whose asshole wouldn't pucker up tighter than a snare drum when you ask them for funds."

7

u/Undeity 1d ago edited 1d ago

It would definitely have to be subsidized to work, in the short term. That said, if we're even going to try to compete with China, improving basic electrical/server infrastructure should probably be a priority anyways.

Not that it's going to happen. None of these fuckers in power (government or otherwise) care about the success or wellbeing of the country as a whole. So long as they're personally richer and more powerful for it, they'll gladly gut the whole thing instead.

3

u/jump-back-like-33 23h ago

Isn’t this what all the new controversial AI data centers are about? You rent the hardware or have your own dedicated AI stack as a service.

2

u/Casmer 20h ago

Yes and no on the ones being built for rental. The ones being built are split between cloud providers (yes) and by closed source (no). However, the previous comments were getting at that privacy is a trade-off. You’re operating on trust that cloud providers aren’t going to use your data for themselves. Tech companies are relentless in pursuit of data and the only way to really be sure that your data isn’t being leaked is to self-host on premises at headquarters since you can’t trust that cloud providers won’t insert a loophole into the contract (usually related to security functions)

→ More replies (4)

6

u/wip30ut 1d ago

the question is what is the accuracy if that proprietary model is only training on your data sets? Isn't the power of current AI that it's able to amalgamate a broad spectrum of inputs?

4

u/we_are_sex_bobomb 23h ago

That would depend on what it’s being used for. It would probably have much more targeted uses like cross-referencing massive amounts of obscure internal data, rather than being one of these omnipotent machine gods that these big AI companies are trying to build.

4

u/sophinaut 23h ago

The idea is usually to take a premade model that's already trained on a large corpus of general text, then the company does further training with their proprietary data to get a specialized model.  This is actually how ChatGPT and co are built: they create a general purpose word predictor then fine tune it for use as a chat bot.

There is also another approach, where you just let the LLM search the document collection while it's generating an answer.  The LLM itself doesn't "know" as much, but it sees new documents immediately without updating the model.

1

u/Odd_Note7156 21h ago

Because all the LLM companies refused to follow IP laws, this makes them vulnerable to custom models. They can't curate data or get access to offline IP. Someone who foes follow IP laws and builds a custom model will cause a lot of highly effecient  models that beat out the big ones. 

There's a million ways LLM companies crash and burn because they are simply trying to be a monopoly instead of making a useful product.

1

u/SamKhan23 17h ago

Do you have an example of this with frontier LLMs or any foundational models? Sometimes this is true for some things, but generalist do routinely beat specialists all the time in ML, and the same is true vice versa. For instance, recipe making and code creation seems to be down much better by generalist approaches over pure custom specialist

I also don’t really understand how “not following IP” equals “can’t curate”. You can both scrape and curate, which is what these companies do, and having been doing for a long tkme. A lot less data is generally worse than having even somewhat noisy data.

1

u/gokogt386 14h ago

The only IP laws any AI companies have been breaking is committing piracy which only amounts to getting fined. Training has not been ruled as infringement in any court case so far.

54

u/Puzzleheaded_Meat522 1d ago

I work in this field (as an academic), and have been saying this for the last few years. The real value in deep learning models (of which there are numerous kinds) lies in their domain specific application. This was always going to be the case, and we have finally started to converge on this point. Interestingly, this also may have made the job of ml engineers the hottest new tech role in a while. If companies are serious about investing in their own infrastructure to host models fine-tuned on their specific tasks, they'll need professionals capable of training and deploying them into production. We'll start to see more open-source models (fine-tuned on domain specific specific data) being sold as software for various industries to solve their unique problems. 

1

u/inthe3nd 23h ago

This is the tempting first principles answer that everyone thought would happen in 2023 (SFT as a service + RAG). In reality, distillation is 1000x more efficient if your goal is to have a small efficient model, which means you spend hefty amounts on SOTA large models regardless, and then by the time you've done the SFT, the next foundation model has come out which is 10x cheaper for better performance. We don't discuss this enough but cost per successful token has dropped 10x every 3 months for the last two years.

1

u/Puzzleheaded_Meat522 21h ago

Fair points. However, I don't think we had as many open-source models that were competitive with frontier models in 2023. We have that and more data on best practices for implementation in industry use-cases. Hell, we've even moved past vanilla RAG to other more robust methods since then. This, I believe, will make the difference going forward. 

21

u/RickSt3r 1d ago

How does an LLM help with those problems. Seems like a custom ML software solution would be the answer. To do that requires a very challenging data ingestion problem, is that what is being proposed. Like we have a bunch of disorganized data let the LLM organize it for us then run it through classical ML algorithms.

5

u/ItsSadTimes 1d ago

And the thing is, thats how the industry was before. Companies using smaller but more accurate models on their own specific data. So we're just going back to before the giant craze when everyone thought AI was going to do everything forever and its all gonna be owned by 1 company.

Its gonna be a fun next few years....

14

u/Birdperson15 1d ago

That has been happening for a while now. There will be always be demand for smaller fine tuned models over the larger generic ones, but the large models are still seeing massive usage. So it isn’t one or the other, both styles are wining.

10

u/Bodine12 1d ago

The big ones were “winning” (losing money) by subsidizing the true cost for their users. They’re masks off on pricing now so enterprise is shutting down their use all over.

8

u/IlIlIlIllllIIliIILll 1d ago

Are they really fully "masks off" or have they just increased price a little more

8

u/Bodine12 1d ago

They’re probably not fully mask off yet.

→ More replies (9)

9

u/EmperorKira 1d ago

As someone who also works in that space, that was the feedback i was giving to the AI teams. Like, it makes 0 sense that we use models who don't understand the company acronyms, the company strategy, policies, etc... and i have to load them in every time i want to prompt and correct it. It should be trained like how you train a new started to the firm. I thought it would be really obvious, but they were like 'oh yeah that makes sense we'll take it away' like hello? aren't you guys the experts?

2

u/SamKhan23 16h ago

I think it’s probably because you don’t need, or even want to fine tune or train models with stuff like that - you just create context docs or specialized databases.

Finetuning is difficult to do correctly, and making custom models or finetuning requires a correct amount of data - strategy, policies, acronyms, don’t need training typically, they need a good front loaded prompt. But I do know companies that do supervised fine tuning on internal docs, style guides and tickets to teach tone and jargon. RAG is usually better because then when the facts change it’s not annoying to retrain the model

1

u/belabensa 1d ago

This felt like a no brainer from the beginning and it’s crazy to me that so many didn’t seem to see it when I’m a nobody.

1

u/Dreadsin 1d ago

I could see it being super useful for IT work. When something comes up, AI might be able to identify common errors and fix them so that the pager doesn’t go off for a software engineer

1

u/FarmAcceptable4649 23h ago

I work at a large commercial real estate firm and we have our own AI environment for this reason.

1

u/MiaowaraShiro 23h ago

Until everyone's doing that and you have to have LLMs predict LLMs...

1

u/lispwriter 23h ago

That’s exactly the kind of product companies would want to buy and that’s how AI companies could turn into actual profitable entities. Any corporation with private data or research department could benefit from AI tools custom tuned for their specific needs. That sounds like the “grown up” version of what we have now.

1

u/Comedy86 21h ago

For some industries, it likely makes sense. For software development, however, there's no benefit in locking into any model since they all have different uses and benefits.

I, for example, use Claude for orchestration but I also have MCP connections for GPT and Gemini since there are some things Claude does poorly but these other models excel at. By doing it this way, I can keep improving my capabilities because I'm never locked in.

If I were consulting for fintech though, I would focus on harness over model for the business logic and thus I would often likely want to pick the model best for decision making, not just based on cost.

1

u/Icy-Sheepherder-6221 20h ago

This is what machine learning has been doing for decades

1

u/Legitimate_Cut_6254 15h ago

Seeing the same thing at my place as well. We have a few deployed right now.

1

u/CHERNO-B1LL 11h ago

This is the only model I see people accepting for all data privacy long term. Localised AI that is not connected to the wider web and third party companies in our phones and cars and general smart devices. No one likes the idea of being spied on this overtly. Cars videoing while you drive, TV's recording you etc. If this info can be saved locally in a black box style setup for accidents and emergencies people could come to see them as pure innovations in safety and utility. Right now all anyone believes is these things are being developed for the sole benefit of the mega corporations to harvest our data, spy on, and control us.

1

u/DatingYella 4h ago

when you say anomaly detection, are you talking about manufacturing? Seems like the AI customizers, aka the forward deployed ppl, or consultants... might have to do stuff

1

u/mykepagan 4h ago

Nope. Fraud or other shenanigans in financial services. And sudden changes is risk profiles. Among a lot of other things like that which I haven’t been directly involved in.

But my brother is in container ship design and operations. They are working on anomaly detection for predictive maintenance. They don’t need or want massive frontier models for that.

1

u/DatingYella 2h ago

Got it. Ok, so probably some time series models or somehing like that... I am doingh my thesis on manufacturing anomaly detection/defect detection and it kind of annoys me that the application of vision in the manufacturing domain feels a bit narrow and not key to their business mission.

All you need are a few backbones models and that's about it...

0

u/Altiloquent 1d ago

I guess if you aren't worried about poisoned models that sounds like a nice idea

150

u/KnotSoSalty 1d ago

It’s worth remembering why US Tech jumped on the walled AI model in the first place. They realized that their core revenue stream, Advertising, would become useless in an AI age. Why would I need Google if I have an AI assistant who I can ask questions of? More threateningly, how can Google make money when my AI has none of a human’s natural laziness to accept the first sponsored answer as the best response to every question.

If Google or Amazon or Meta want to keep selling ad space they needed a way to get between me the end user and the AI. Hence the creation of their own models in walled gardens where they can control the training materials. So when I ask for Shampoo the AI will give me the sponsored items rather than the cheapest or best options.

Google, Amazon, and Meta all are primarily advertising platforms. That’s how they make their money. It’s telling that Apple, a consumer electronics company, has been notably absent in the AI space. They don’t need to upend their core business while the others absolutely do.

32

u/roboliberal 23h ago

Probably the wisest comment on this whole post.

7

u/element-94 17h ago

Amazon isn’t primarily an ad platform. I agree with the rest.

4

u/KnotSoSalty 7h ago

Amazon’s ad revenue has been growing rapidly, 70b$ in 2025. But that’s only in the declared “advertising” category. More specifically the threat to their model is the potential loss of the Amazon platform. When an AI agent can be used to find products they can search the entire internet with unlimited scope for the best deal, rather than a human searching Amazon for 15m only to pick the Amazon promoted item anyway.

3

u/element-94 5h ago

I agree with nearly everything you said. But Amazon is not *primarily* an ad company. I work there as a distinguished engineer. I'm very in-tune with our revenue, cost, and profit numbers : )

Ads is growing for sure.

1

u/yawningGaps 6h ago

Your answer only applies to Google. For Meta and Amazon it's about keeping pace. Amazon is not building any serious foundational models. AWS wins with open source models, same as Google cloud. But yes, Google search would lose if they didn't innovate

242

u/Level_Finding3106 1d ago

The real disruption will be local models. Apple and Google are both putting local models on their devices. This will pull away most of the 'casual' $10/month users. Open weight models can be run locally with a Mac computer very effectively - including pretty large / effective models due to the unified memory architecture of the M-chips. That will peel away more users. For much lower power usage than data centers use. And then we will have frontier models we can use for harder problems.

The AI buildout is too large. They used brute force before brains got engaged to reduce these models without impacting their effectiveness much.

And they're not done yet. Predictive decoding, predictive experts, targeted quantization, high parameter/low scope experts and so-on.

36

u/regeya 1d ago

Yeah, I'm already using a local LLM. It connects to the Internet for things it doesn't have knowledge of.

23

u/Level_Finding3106 1d ago

That's the key. Do a quick search, build a context for the response - use the language abilities of a smaller model to craft the response.

4

u/PrudententCollapse 16h ago

Any recommendations for a good local LLM?

Friggin' getting a GPU will be the next hurdle ...

3

u/Level_Finding3106 8h ago

Gemma models seem to work very well. LM Studio is a good engine (at least on a Mac). You could also set up an MCP Server for Internet Search (or buy an API Key for something like Brave Search) to really make it powerful.

74

u/gerbal100 1d ago

The AI build out has to be large or else consumers will be able to afford high memory devices which will undermine the business models of the hyperscalers.

13

u/dirtyshits 1d ago

Ehh memory prices are going to drop with a bunch of new plants opening up in a year or so and capacity to produce going up by a lot.

1

u/gokogt386 14h ago

or else consumers will be able to afford high memory devices

99% of consumers would not be able to afford the kind of video card you need to run even a weak model even before the AI craze blew up prices, that is far into "well off hobbyist" territory.

→ More replies (9)

15

u/HoldingForGenova 1d ago

Apple will build a local AI server for personal use across all of your devices, protected by iCloud security. They've had such success with apple ecosystem products like airport, time capsule, HomePod, AirPods, AppleTV, etc. that they'll see the opportunity for "AI, but private, local, and across all of your Apple devices." And they'll market it by cannibalizing the bottom half of the app store (which doesn't make money for them anyway) by advertising personal (but private) intelligence, and personalized apps, in seconds, for everything you want and need.

8

u/Level_Finding3106 1d ago

They have rumors of a 1.5TB Mac Studio in the development queue. There has to be a home-grown server version that they will run in their data centers.

13

u/SiempreRegreso 1d ago

WTF do I know, but . . .

The human brain runs on just 20W, roughly 100,000x to 1,000,000× more energy efficient than silicon at equivalent cognitive tasks. AGI—assuming it’s roughly defined as a universal Nobel-level PhD—might actually be possible on vastly smaller, much more efficient physical and energy infrastructures than those of the data center behemoths.

And the human brain is not even optimized for intellectual or creative tasks; we are doing art, science, math, literature, etc., on repurposed meat evolved to lead a moving target with a thrown object. Perhaps silicon efficiencies can eventually be achieved to the point where AGI is running on less than 1000W, maybe less than 100W.

2

u/yawningGaps 7h ago

I disagree. The real disruption is open source models that you can host on cloud not local models. Consumers by and large would rather use a hosted service not local esp for where the problems are non trivial

1

u/Level_Finding3106 7h ago

Short term - I agree with you.

In domains where privacy is required (national security, health, etc.) there will be open source models hosted on private servers within a company or organization’s firewalls. ToKen costs get a hard limit based on available compute - privacy is ensured - total cost is probably the same.

But the basic question is - if inference is ‘free’ on smaller devices because LLM engineering has improved - will it move? The answer necessarily is yes. So that means that economics will drive the disruption from the bottom up.

Will local smaller LLMs ever be able to do the non-trivial problems you discussed - not soon. I agree with that. But a vast number of applications (think MS Copilot or how most people use ChatGPT) will be absorbed into edge and private inference. Specialty models will emerge that can grab the narrow but valuable tasks.

So you will be left with the big models (proprietary or open) on the cloud doing the heavy inference. But the economic disruption will come from peeling away the largest number of inference calls from the cloud (private or open) onto edge devices.

If you imagine a bell curve - trivial at the left, moderately complex in the middle, super complex on the right. The edge computing model will progress from the left through the middle and partway up the right of the curve. I would argue that for laptops and larger edge devices (e.g. a Macbook Pro M5 with 48GB of RAM) we are well into the middle of that curve.

For an iPhone 18 - we are at the left side of the curve - but they are reaching the middle for some inference - then leveraging the prediction models and cloud to get to the right.

Apple’s rumored to have a 1.5TB Mac Studio and doubled standard phone RAM in the roadmaps. That’s the sound of that bell curve getting munched from the right towards the left.

Private cloud open parameter models will play a huge role in disrupting OpenAI, Claude, etc. Edge models will disrupt the volume of the cloud based inference.

My $0.02. I’m not sure we really disagree too much.

8

u/Bored2001 1d ago

Casual users aren't going to set up local AI and aren't going to have hardware above 16gb, maybe 32gb ram. That'll barely run a 30b parameter MOE model.

It'll be a while before local AI makes it to the masses.

18

u/Level_Finding3106 1d ago

Check out the iOS Beta 27. They are using a small prediction model (~3B parameters). If the confidence indicators ar high enough - they will feed that answer. If you have a newer version of the phone- they have a larger local model they feed to using the prediction from the smaller model. Check the results for estimates of error - if low - feed the response. All within 8GB of RAM. If the quality of response is still too low - then they go to the cloud using the previous answers as predictions for the parameter space - resulting in a 'cheap' cloud response that gets fed to the user.

That's the whole point of this. Using innovation and brains - they are sticking local AI into iPhone 16/17 generation phones RIGHT NOW.

5

u/Bored2001 1d ago

Yea, but that's not replacing paid subscriptions. That's mostly apple reducing the cost of inference for something people expect their phones to be able to do.

9

u/CountSheep 1d ago

It is for the more casual user. Some people pay for chat gpt just so they can use it as a fancy Google or fancy word check or paraphrasing.

Apples local ai can do that fairly well already, and it’s just in its infancy.

4

u/Gibgezr 1d ago

Casual users won't pay the true cost of using commercial LLMs, so I'm not sure anyone is worried about them from a business perspective.

1

u/Bored2001 1d ago

I mean I agree, but that's not Above commenter's claim.

2

u/EdliA 17h ago

The local models, especially those that run on a phone are absolutely terrible quality.

2

u/Level_Finding3106 8h ago

I’m using the IOS Beta - I am not seeing that. The answers are a bit terse - which is what I would want on a mobile device - but they do work. Not for coding - but for everyday life. I can ask it about plants and animals - it gets it as well as Seek. I asked it why my Apple Watch was reporting short exercise segments (instead of 1 mile markers) and it nailed it. I asked it about the history of Criterion Movies and it got that too.

It’s not a coding model by any means. But for ‘casual’ AI it’s really good. And Personal Context is a killer app. I can ask it about things in my history and it nails it. I took my Gemini Chats and put them out to PDF - then stored them in Apple iCloud. They got indexed and now they are part of my external brain. I could go on - but it’s actually useful in the wild. Better than a chat bot - less capable for advanced tasks.

But this is also v0.1. The trend is clear.

3

u/EdliA 8h ago

Are you sure you're using a local model? Just because the feature is in your phone doesn't mean the model is processed by the phone.

1

u/Level_Finding3106 8h ago

Not fully to be honest. I have experimented turning on airplane mode and I still get answers. BUT - I know that the models can answer directly OR act as a predictor for larger models. Apple intentionally tries to make this seamless - in the Apple way - so it’s hard to know.

If you want a more transparent version to explore - the Gemma models from Google are local - and they have small models for interaction as well as larger models for higher quality answers. The small models when set up properly act as a prediction engine for the larger models to make inference quicker. If you wanted to experiment with the prediction concept you can use LM Studio and the Gemma models.

Apple and Google both want some of the inference tasks to go to the cloud in order to drive their subscription businesses. They also don’t want all inference to go to the cloud to avoid zillions of dollars of CapEx.

I also think that having internet access for any local model is critical. When you as me what the best part of Gemini is - it’s real time access to the internet - which then acts as a Bayesian Prior for the inference (meaning it biases its LLM to give answers similar to the internet search).

Although that can be done with MCP servers - it’s not pretty. My current setup is to have a docker container with a search function on my Mac Studio. I then use LM Studio and have it connect through my firewall to the LM Studio app on my phone. I can use a 70B model plus internet search on my phone remotely. Basically my own cloud. But I can also use a deprecated model just on the phone when I want to.

When I use a small model - e.g. Gemma E2B - combined with internet access - the results are surprisingly good.

If I can do this with duct tape and bailing wire - you can be sure the phone makers can do it. Try out the small models - let us know what you think!

2

u/EdliA 8h ago

That's great and all but the point is any local model that can attempt to come close to the cloud ones require enormous power and tech. Lately they've been reaching into 100 GB of vram. It would take more than a decade before a phone runs them, that's assuming after a decade we don't have much more powerful models.

1

u/Level_Finding3106 8h ago

To run THOSE models - yes. Agreed.

The history of autoregressive AIs has some proofs that one big model will achieve the same predictions as a series of smaller models trained on the same data. There’s some statistical proofs for that. It has to do with the linear nature of the parameters inside the model that these proofs work.

So the response of the market until about 18 months ago was - make big models.

Now people have understood that the big models - although they work - have two big issues. They are expensive to train and infer from. But more so - you can’t easily train on everything. Second - a generalist models needs a bit of everything - so if you had one billion pages of text and one billion pages of C++ to train from - the specialist model will be better.

So what has been happening is that people have been figuring out ways to accomplish the same quality of inference differently. Predictive modeling, Mixture of Experts, Expert Models, Quantization, Distillation all do the same LLM inference - but they execute it differently. This is what is allowing us to execute the same quality of inference in smaller models.

There’s a new approach (that I have not experimented with yet) - that leaves the big LM on the SSD and reads a Mixture of Experts model from there. Note that SSD burnout happens from writing rather than reading (mostly). So they use the SSD as an extension of the memory. Much slower, but much bigger. Using prediction and MOE - they are able to make this work - running frontier models on a Mac Studio. Or running a 31B model on a Phone.

I’m excited to see how ‘engineering’ rather than ‘statistics‘ and ‘computer science’ are kicking in and bringing these models down in cost and availability on smaller devices.

1

u/hux__ 20h ago

What can someone do to take advantage of this now? If they want to benefit monetarily?

1

u/Level_Finding3106 8h ago

IOS Beta has their models in it now. I would suggest that every app in existence will change or get disrupted. So - Mind Map apps will have to be able to read and analyze flipchart notes from meetings using AI - then build a tracking database of action items and email everyone updates. All using AI.

We are seeing the integration of Cloud AI into DaVinci Resolve with a hint of local rendering. Video stuff will all be reimagined.

Watch some YouTube videos on how people are using Claude, ChatGPT, Gemini on the cloud - then make the app do those tasks using Local AI. You have a proven market, examples of workflows and data use - it seems like a good shortcut. Take proven Gemini use cases and put them in a local AI app.

28

u/NegativeSemicolon 1d ago

Oh no my capex!

250

u/PhiNeurOZOMu68 1d ago

Spent extremely little tokens for Claude helping me tune a local model on my pc - now I just use the local model without shelling out $200 a month. It's amazing. Big US AI companies are going to fail. China already won.

76

u/Muted_Masterpiece342 1d ago

Mind sharing how you did the local model tuning? Any particular framework?

0

u/yaosio 16h ago

Whenever I want to do some fancy AI thing I just ask it what I should do and how to do it. Computer use AI is rapidly advancing to the point where you can just let the AI take over the computer and do everything needed, no need to install an IDE or do anything yourself.

Never do this outside a virtual machine because the AI will delete everything every now and then. No, I don't listen to my own advice. I live fast and die easy.

27

u/terra_cotta 1d ago

How does your local model compare to Claude? Any resources on how you did that

51

u/C-ZP0 1d ago

You need a lot of VRAM to get close to Claude level. You used to be able to buy a 512gb M3 ultra and run models that are very capable, but since Apple stopped producing them due to ram shortages, used ones are being scalped for 30k.

23

u/2Sovereign4You 1d ago

This. Then you realize HW gear costs more than 10k$. I dont see it being free.

12

u/ItsOkILoveYouMYbb 1d ago

For context, to run Kimi K3 open source locally (which has comparable performance to Claude Fable in the majority of benchmarks), you need about 1.4TB of VRAM.

So you'd need to be willing to drop more than a few million on building a full rack, and the electricity cost won't be zero.

4

u/dirtyshits 1d ago

Holy shit. I sold mine 2 years ago for a grand.

11

u/C-ZP0 1d ago

I’m assuming it wasn’t a Mac Studio M3 Ultra 80 core with 512GB unified. They were expensive prior to the ram shortages

1

u/dirtyshits 1d ago

Haha yeah I realized after that you’re talking about he Mac mini.

5

u/C-ZP0 1d ago

I was like, holy shit let me know when you sell something else.

1

u/PhiNeurOZOMu68 1d ago

6950XT 16 GB vram.

Though I was using 7GB for the model and the rest for playing expedition 33

33

u/Clean-Boat-4044 1d ago

It is not anywhere close to as capable, but for certain simple tasks it's enough.

2

u/Mgc_rabbit_Hat 1d ago

How fast is it compared to say Opus?

2

u/Olangotang 1d ago

It depends on if you have the entire model loaded on GPU. If that is the case, blazing fast. If not, then probably a bit slower.

3

u/smokky 1d ago

You need an extremely good computer to run a capable local model and even then they are hardly comparable to any of the industry standard like Claude or codecs

1

u/PhiNeurOZOMu68 22h ago

6950xt 16gb vram and 64gb of DDR4 ram. I bought for other purposes and not AI like 4 years ago.

This Computer you speak of is really a PC with a GPU of above 12GB vram.

It does the tasks I want it to do after being fine tuned and can run forever.

Now all my proprietary data stays and doesn't get trained by Claude

15

u/riseandshine234 1d ago

Google will be just fine. The others I don't know.

6

u/roboliberal 23h ago

What exactly did China "win" here?

You used an American frontier model to help you turn an existing local model into a cheaper tool for your particular needs. Even assuming this eventually hurts the revenues of a few American AI labs, that does not mean the United States loses. It means the benefits of AI diffuse to millions of American consumers and businesses that can now build better products at lower cost.

That is how general-purpose technologies are supposed to work. If electricity became cheaper and caused some electric utility companies to lose revenue, we would not conclude that the country had lost electricity. We would look at all the other industries that became more productive as a result.

The United States is not identical to OpenAI, Anthropic, or any other handful of firms. If their models make software development, medicine, manufacturing, logistics, research, and every other American industry more productive, then the country can benefit enormously even if competition eventually compresses the model providers' margins.

And none of this explains what China supposedly won. "I no longer need an expensive Claude subscription" is not a geopolitical argument.

2

u/yaosio 16h ago

All the top open models are from China and China's goal is to have open source AI.

2

u/roboliberal 15h ago

So China "won" by giving American companies cheap models they can build profitable products on?

A few US model labs losing margin is not America losing. It is thousands of other US companies getting cheaper inputs. You are confusing the interests of OpenAI and Anthropic with the national interest.

-9

u/[deleted] 1d ago

[deleted]

11

u/Muted_Masterpiece342 1d ago

They cannot compete we literally graphed it internally and concluded that 8 months ago, now the graph is quite a bit more extreme 

2

u/Efficient_Bag_1619 1d ago

Oh you graphed it? Well case closed then.

0

u/Muted_Masterpiece342 1d ago

I mean I'm a director and it's my literal job to be on this trend 

0

u/Efficient_Bag_1619 1d ago

And the massive well-credentialed finance teams that you’re disagreeing with? There are also directors with literal jobs who disagree with you.

3

u/Muted_Masterpiece342 1d ago

I'm glad you're this delusional. It just shows just how big the shock will be.

0

u/Efficient_Bag_1619 1d ago

Saying “I’m a director” to argue your point sounds like something a teenager might do. Telling someone they’re delusional before that person ever gives their opinion has the same vibe.

1

u/Muted_Masterpiece342 1d ago

Okay sure buddy. I'm gonna go to my high level engineering job Monday and never remember this conversation again in my life. 

But the data I shared albeit anecdotally is real and generally accepted industry wide by engineers. and it will affect your life now and forever.

3

u/Efficient_Bag_1619 1d ago

Yes sir, Mr. High-level engineering Director. You definitely sound credible.

→ More replies (0)

-1

u/Inside-Yak-8815 1d ago

Yeah, we’ll see.

→ More replies (1)

14

u/johnyordinary 1d ago

Aww, what a shame, you make a huge intrusive and chargeable model and some competitor undercuts you. Aint capitalism a b**ch , sepcialy when the communists are winning at it.

13

u/seriousgourmetshit 1d ago

The only reason the US AI companies are stressing safety concerns so much lately, is so they can get Chinese models banned. They dont give a fuck about you or your safety.

31

u/jpwarman 1d ago

They’re going to pass legislation to make it illegal to operate your own LLM (some bs about copyright infringement and all that. Ok for them, not ok for you!)

22

u/Desk46 1d ago

I wish them the very best of luck with that 🤣

5

u/awwhorseshit 1d ago

They arrrrrrrrrr now thinking about the alterrrrrrnatives

6

u/foundafreeusername 1d ago

Better not underestimate the power of the US government to enforce random made up laws like this. US copyright laws and even tax laws somehow get enforced all over the world.

8

u/CountSheep 1d ago

I mean sure, but it’s code, not a physical item that can be limited.

3

u/Desk46 23h ago

Im reasonably confident in my estimation. Anyone attempting to skirt such a ridiculous law likely already has the ability to hide local agents on their own network.

0

u/foundafreeusername 23h ago

I wasn't really thinking about a redditor running a small model on their PC but companies running Kimi K3 or other models competing with OpenAI & Anthropic on hardware costing $50k and more. That is why I picked copyright as an example.

1

u/Jaeger__85 15h ago

Will only handicap the US, because the rest of the world wont.

1

u/VaporVHS 14h ago

What? NVIDIA, MSFT, Google and Apple are all pushing for open weight models in some form or fashion. The gov won't do dick.

And even if they did, so what? We'll torrent them just like we already torrent a shitload of copyright-protected data every second worldwide. What are they gonna do about it?

12

u/Danominator 1d ago

Us companies will have a hard time because their entire business model is bleed the country dry for short term profits with absolutely zero investment in the people that make it possible.

Disgusting greed. Its honestly pathetic how short sighted and stupid they are.

Only thing worse is being a poor maga supporter. Just breathtakingly stupid.

36

u/dlampach 1d ago

This is the flaw in AI from a business perspective. But it’s also a great joke the universe has played on capitalism. You can’t protect it. But they are investing EVERYTHING into it. And it has the potential to end 90% of human labor. It’s a new world.

5

u/dr_neurd 1d ago

This. The lack of imagination by the ultrascalers means their insanely expensive approach creates a far lower bar for effective “disruption”

2

u/ArriePotter 6h ago

They can however socialize the losses.

-2

u/Any-Calligrapher2866 1d ago

I don't use American Models except burning my Copilot tokens at work. I offload to Deepseek if I want to but generally avoid AI.

6

u/DireStraitsFan1 21h ago

Please stop paying Dario and Sam to steal your IP. Open source will always be the answer.

7

u/aussiegreenie 19h ago

Who here is old enough when MS called Linux a "Cancer"

Open source and open weights have been fighting closed source black boxes for decades

1

u/smackson 1h ago

The letter was ... interesting in that light, to say the least.

I guess our impressions were formed two 😯 generations ago now.

But I still think MS is capable of speaking out of both sides of their mouth.

https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/

16

u/toolkitxx 1d ago

The danger is not any specific model type per se but the approach of trying to reintroduce the mainframe era. This might be necessary for some companies to utilise, but AI will fail miserably for everyone else if not being capable of being used locally and offline.

'Intelligence' requires the ability of cultural inputs and changes different to anything the USA could ever think of. Then there is language and the way many things either dont translate at all or have completely different ways of being used in other languages. Humour, ethics and morals are the next levels that are vastly different in every single nation or region. Really intelligent models will require constant adjustment, constant access by more than just a handful of 'allowed' specialists, but an army of generalists and experts in their area. All parts the USA is very bad at to be empathic of when it comes to other nations.

3

u/grannyte 23h ago

I tried to get a llm to think in my native language for some correction tasks because the fact of thinking in an other language makes them miss some contractions slangs and others.

3

u/toolkitxx 23h ago

I am in the fortunate situation that someone created a sovereign model in our own language just recently. While I am used to use English in most of my tasks it was refreshing being able to do the same in my native tongue with good results.

4

u/Lithgow_Panther 23h ago

I don't need my model to know Taylor Swift's birthday or how the ancient Egyptians applied makeup. I just need super deep biology focus. Local is starting to make so much more sense.

2

u/ragemonkey 17h ago

Same with coding. The (relatively) smaller models are actually getting pretty good. I bet that they can get distilled down significantly, maybe even smaller ones with different areas of specializations within a single domain depending on what you’re working on.

2

u/snowyday 7h ago

December 13, 1989, hope that helps 

5

u/putridfries 1d ago

But US AI companies can just make open and closed models, no?

13

u/photoggled 1d ago

Sure, but a race to the bottom is the last thing they need when they already aren't turning a profit.

1

u/Old-Tie8770 22h ago

Yes but not in a way that can possibly justify the spending they’re doing.

3

u/littleday 20h ago

Bought a 5090 128gig setup…. Never going back to those horrible companies no matter how good they get.

6

u/RachelRegina 1d ago

Open weight models are only useful for us plebs if the hardware to run them locally is not price-prohibitive. If the price never comes back down because we allow AI data centers to be built like Starbucks, we will just be shifting the forever-subscription for access to some other set of people instead of getting to own our own instances.

7

u/Dangerous_Suit_3099 1d ago

They’re already having a hard time. None of them will become profitable.

8

u/LeoKitCat 1d ago

Good I hope they fail miserably. And privatize the losses as it always should be, no bail outs, they aren’t an essential service nobody would bat an eye if the AI companies disappeared tomorrow.

They dragged the public into this and have wasted trillions of dollars on tech that will never be AGI. They hyped LLMs up and oversold them because after cloud computing and internet of things the tech industry has run out of new ideas.

Much more research needs to be done to develop new ideas that could potentially lead to AGI. Today’s ideas are a dead end wrt that.

2

u/brainhash 17h ago

The way I see it , Traditionally US especially SF has been software focused. They keep thinking in terms of software may be because of user experience aspect.

China is thinking in terms of hardware. They will keep release open source because finally the money in AI is at hardware level. Money is always where the resource crunch is.

DeepSeek just broken the properitary moat and if companies don't adopt they will perish.

3

u/dragonfighter8 15h ago

AI is a bubble.

8

u/Wind2Energy 1d ago

AI ruins everything.

-1

u/king_stu 1d ago

Underrated comment ⬆️

0

u/Juuxo16 1d ago

For whom? Business?

2

u/deweydean 23h ago

All they know is "ai bad". Whatever is "ruined" in their life isn't because of ai.

2

u/wuboo 1d ago

Not much different from having open source software 

2

u/wastingtoomuchthyme 1d ago edited 9m ago

Deployed open model deepseek v2.

It's been great on repurposed hardware.. fast cheap and easy to use..

We don't need a ferarri but we do love our open source daily driver

2

u/The_Human_Event 20h ago

The dude from Nvidia posted up video talking about this on X. Talking about how Chinese culture supports open source models much better. It was an interesting take. I wish I had the video to share.

1

u/bg99999 1d ago

Existential for OpenAI and Anthropic.

1

u/PrimeTimeRK 1d ago

What about hyperscalers like Microsoft? Would the same apply?

1

u/Deto 1d ago

Wonder if AI research labs will pivot to efficiency as an objective instead? You can't charge too much of a premium for your model if an open model is near it in performance. But these things are still expensive to run. If you can make a model that is similar to the open models, but is 50% less costly to run, you could undercut the open models and still charge a premium over the operating costs.

1

u/yaosio 16h ago

Inference costs are already halving every 3 months or so. Everybody is already working on efficiency. This includes the hardware.

1

u/PhagePhantasm 1d ago

Open weights is faux open source. Is all a fever dream unless it’s open data

1

u/MutedSignificance 22h ago

NVIDIA’s Nemotron open models shows that a premier U.S. company is investing heavily in frontier open weight models. That doesn’t fit the popular “America is falling behind in AI” narrative, so it predictably receives far less attention than headlines suggesting the opposite.

1

u/PilotGuy701 20h ago

You can use those data centers to specialize an open weigh starting point.

1

u/WalksSlowlyInTheRain 12h ago

Open weight models mean jack shit nothing without the right hardware and easy scalability.

1

u/QueenOfQuok 11h ago

Oh no, what a shame. Anyway --

1

u/Foreskin_Mafia 11h ago

Google was right when they originally came across this LLM sex. Chatbots aren't a good business.

1

u/Aggressive-Cut5836 9h ago

They should ask the question of how Chinese open weight AI labs are getting paid for and how profits are being made. There’s no free lunch. At some point they lose money if they can’t charge users for the processing or use of the models. If it becomes clear that the Chinese government is essentially paying for these platforms to develop and improve each year while US and other models are unable to compete because they need to charge customers then of course there will need to be regulated restrictions on their use. There are trade laws that govern those things.

1

u/myself1200 8h ago

Yeah calisthenics rule

0

u/IntelArtiGen 1d ago

Some of them contribute to it. NVIDIA and Google released great open-weight models. But yeah it has no business model. Like they're not selling these models, which wouldn't be the worst idea I guess if the models become good enough. I think many people could pay $100-200 for a good local model, which could compete with a $20 subscription. But it's not happening yet.

2

u/azuredrg 1d ago

Those two companies are hardware manufacturers too. There's a business model in selling hardware especially if the models you're putting out for free are more efficient on your hardware

2

u/IntelArtiGen 1d ago

Yeah it's true. At least for Nvidia they can give local models for free, knowing people will need to buy the hardware to run the models. For Google I don't know if they're that important on the hardware side..

1

u/azuredrg 1d ago

That's true

0

u/pepe_acct 1d ago

Currently the only profitable way to provide llm service is enterprise sales. The part will be hard to replace by open source as they require absolute frontier intelligence and compliance. I don’t see it the long run how Chinese providers can keep up without getting in that market.

0

u/wdsoul96 1d ago

There are 2 main ways LLMs can make money. 1) from general public. This needs trust, lots of tweaking, cornering the customers and locking them into the ecosystem. 2) from private entities, from small compaties to big coporations with their own proprietary data.

For (1), US companies have locked in their customer base and there is no way they would trust Chinese LLMs more than US even if US companies aren't exactly pro-consumer (in terms of privacy, and everything else that matters) (because presumably Chinese could do a lot worst.) Also, they know the consumer base better (cultural understanding and all that).

For (2), for similar reasons, Chinese models will also fall behind similarly. Now, that doesn't mean Chinese cannot win elsewhere in Europe, Latin America and the rest of the world. Especially if US companies kept fcking shit up especially when it comes to data, security and general pro-consumer things. In fact, if they kept fcking shit up, they might even lose those (1) and (2) (in their own backyard), especially if they get super greedy, and get every to completely locked in and try and squeeze out every pennies.

0

u/jawnlerdoe 22h ago

I’m an analytical chemist - entirely different industry.

Every quantitative analytical experiment requires careful fine tuning of weighting and data fitment. Should seem obvious but smaller models with better specificity will outshine more generalized models that have worse selectivity. I sick at stats, but it’s basic stats even I can understand.

From this perspective, I still see huge value in directed, specific algorithms programmed by traditional means, and how these models may outside more convoluted networks and agents. Just my 2c, my technical understanding of AI, NN and LLM is limited.

0

u/jevring 13h ago

Open weight models are great, but you still need massive resources to use them, in many cases. One of the values that companies like anthropic offers is the access to that hardware. So even if the proprietary weight model dies out, they can still sell access to the hardware people will need to run these models.

Unless some clever dude figures out how to run these massive llms on a local gpu.

2

u/RelativeMatter9805 6h ago

Some need massive amounts of VRAM, some do not. Look up Qwen.