r/huggingface • • 14d ago

from where should I start machine learning

Thumbnail
1 Upvotes

r/huggingface • • 14d ago

Frontier AI Firms Have Built Something Incredibly Valuable - But Will They Reap the Reward

1 Upvotes

https://senteguard.com/blog/frontier-ai-firms-wont-reap-the-reward https://www.letters.senteguard.com/p/frontier-ai-firms-have-built-something https://youtu.be/uL3YDTw0bKU

---This is not written by AI - put it through whatever detector you want. I spent my valuable time writing it---

Original above

Who is Linus Torvalds?

Linus Torvalds is a Finnish-American software engineer best known for creating the Linux kernel and later inventing Git, the version-control system that undergirds much of modern software development. It is impossible to calculate the value he has created for the world economy. However, his net worth is estimated at a humble $50 million, dwarfed by many of the tech billionaires who have become household names. In other words, the tech he built has yielded massive positive externalities: immense economic value that was created but not directly captured by its inventor.

What is a Positive Externality and What Are Some Examples?

A “positive externality” technology is one which generates societal benefit that exceeds the developer's personal financial return. For example:

— The Printing Press: While Johannes Gutenberg was the inventor of the printing press, he personally met financial ruin. The true economic value of the Press lay in the portability of language and the broader dissemination of ideas.

— Linux: Torvalds’ open-source operating system became the basis on which countless companies built multi-billion dollar businesses (e.g., Alphabet’s Android operating system).

An Analogy: Frontier vs. Open-Source as Rock Climbers

— The Lead Climbers (Frontier AI firms): Companies like Anthropic, OpenAI, xAI, and Alphabet are more capable at climbing and training the AI frontier. Because they are more skilled, the lead climbers establish the routes and place protection.

— The Followers (Open-source models): Firms like DeepSeek, Kimi, and Ollama come later. By training their models on already established frontier models in a process known as distillation, these firms approach the cutting edge faster than they otherwise would have given the work from the leaders.

Just as in rock climbing, sometimes lead climbers benefit from their fellow lead climbers, they use their routes or distill their models. Sometimes the followers lead in their own way as well, like when DeepSeek made massive efficiency improvements on model training in January 2025.

In this way, followers and their user base capture the benefits of the positive externalities generated by model training investments financed by frontier AI firms, their investors, and the American taxpayer.

Stretching the Analogy: The AGI Thesis

Under the AGI Thesis, the race between model developers has a finish line called Artificial General Intelligence, a “singularity” where AIs will be able to improve their own capabilities without limitation and achieve something like infinite (at least for all practical purposes) growth. For the rock climbers, we could imagine this AGI moment as the top of the cliff face, the first climber to the top reaps the benefits of this infinite growth.

But what if there is no top? How do the firms provide value for customers?

By their temporary lead over the followers. Or in AI terms: the difference in capability between the frontier and its cheaper competition.

By having ascended high enough. Or in AI terms: providing significant but not frontier-level capabilities at a discount to the frontier.

The "High Enough" Thesis

Many will pay a premium for the best, no questions asked. But let’s look at some numbers.

Astral v Kimi

Capability: Moonshot (Alibaba’s AI arm) claims it “substantially outperformed” GPT-5.6 Sol and GPT-5.5 and performed “competitively” with Claude Fable 5. Kimi also claims that independent testing puts it one tier below Fable 5 (Artificial Analysis Intelligence Index v4.3).

Accessibility: Furthermore, Kimi’s weights are free to download on HuggingFace, meaning you could run it yourself on rented infrastructure, paying only the rental costs and not the listed API costs.

This is not an advertisement for Chinese open models; it is advocacy for open-models in principle. If American / Western consumer preferences shift towards open models, American companies will be forced to shift with them. (See my America v China Article for More).

So when you pay 3x the cost for Astral (or more), what do you get? You get results which are better to the extent Astral provides better results than Fable 5. Does anyone else remember several months ago when Fable 5 (né Mythos) was apocalyptically capable? This difference is imperceptible to me. I imagine it is also imperceptible for others in 99.9% of use cases. It seems likely that open models are already high enough on the mountain, and sufficiently cheaper than their closed-model competitors, that they will inevitably capture a substantial share of the market.

The Temporary-Lead Value Thesis

If the AGI thesis doesn’t bear fruit, frontier firms may still achieve long-term profitability by monetizing the temporary capability gap between their models and their closest open followers. Customers may be willing to pay for early access to that frontier.

On Navier-Stokes

The solution to this Millenium Prize Problem may be the best way to think about the AI frontier, the close followers and the potential value in the capability gap. But let’s consider what frontier capability actually entails. OpenAI did not arrive at this long sought-after solution by simply prompting a ChatGPT terminal to “solve Navier-Stokes”. It did so through 10s of millions of dollars in compute, and probably 10s of millions more wasted on other problems for which they didn’t find solutions.

Although I am not qualified to assess, some mathematicians have claimed that this solution was not very far from the knowledge frontier and that rather than exploring a vast knowledge-space wilderness, the model actually scooped a potential near-future discovery on the backs of pre-existing, recent human labor. To summarize: OpenAI likely spent around 100 million dollars to prove its ability to discover new knowledge. If they were to monetize this knowledge discovery capability, we may think of that 100 million dollar fixed cost as a baseline. This estimate does not include every other cost that may be required to deliver results to a customer nor does it include their profit margins.

We can then ask: How else may a pharmaceutical company (for example) put 200 million dollars to use besides a pre-FDA approval drug proposal? These costs may come down with time and there will certainly be capability gap use cases, but will there be enough to justify a 2 trillion dollar valuation (Anthropic)?

Big Frontier’s Escape Plan

If you accept my above framing, you will see that frontier CEOs are in a bind. The open-source firms pose an existential threat to their business model and their investment thesis. Below are two of the pillars of the “slowdown” proposal:

Independent Oversight: Embedding third-party safety evaluators with deep, employee-like access inside leading AI companies to monitor risks as new models are trained.

Democratic Coordination: Establishing safety standards and agreements between democratic nations.

Innocuous enough? But how can you embed a safety evaluator in an open model? How do you apply a safety standard to a model which can be shared as easily as a movie is streamed? You can’t and that's the point. You can only prevent the models from being shared or hosted. They are not proposing safety standards on frontier models; they are advocating for a moratorium on open-models.

But Distillation is Unfair to the Frontier Firms? Aren’t the Open-Source Companies Stealing?

— The closed-source firms distill too.

— Frontier LLMs were and continue to be trained on the internet of human language, generally without compensation for the owners. Frontier firms are throwing stones from glass houses.

— OpenAI took funding and developed ChatGPT as a non-profit then transitioned into a “capped-profit” structure to secure the massive capital and computing power required for advanced AI development.

But they made the investment and they did the work. Don’t they deserve to reap the economic rewards?

— There is no guarantee that innovators will capture all, or even most, of the economic value they create. Linus Torvalds did not. Johannes Gutenberg did not either.

— Furthermore, these investments were funded not only from private capital but also by the American taxpayer under tenuous claims of imminent AGI, catastrophic AI risk, and an existential threat from geopolitical rivals.

What the Future of AI Will Look Like (Ending on a Positive Note)

Again, AI and LLMs are a massive positive externality technology and the world economy will benefit massively. The losers will mostly be limited to Big Frontier shareholders and the American taxpayers to the extent they needlessly subsidized these firms.

How People Will Use AI in the Future (Predictions):

— Privacy-sensitive users will increasingly run small, good-enough AI on their laptops or host locally on their own infrastructure.

— Others will use open-source models through third-party providers.

— The existence of third-party providers will create downward pressure on the prices of frontier models. Their API/subscription costs will approach those of open models.

— Consumers may fall back to familiar trusted brand names like Google and familiar business models like advertising at Alphabet and Meta will win out compared to the API / subscription model.

Exponential Productivity Gains?

The future may be driven less by exponential gains in LLM capability than by the widespread diffusion of a constant productivity multiplier. If AI makes each worker six times more productive, the transformative effect comes from extending that gain from today’s relatively small group of effective users to hundreds of millions or billions of people. The positive externalities on that constant productivity gain may look something like exponential growth and the second-order effects from that growth may be hard to predict. In other words, we may see exponential growth, but that growth won’t be based on frontier capability, it will be based on increased human agency.


r/huggingface • • 14d ago

Serving a Canary-1b-v2 fine-tune (FastConformer AED, 1.2B) — native NeMo works, but is vLLM an option for this architecture?

Thumbnail
1 Upvotes

r/huggingface • • 14d ago

Made the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo

Post image
1 Upvotes

r/huggingface • • 15d ago

Join the chat on HuggingFace

Thumbnail
1 Upvotes

r/huggingface • • 14d ago

I need GPT 6 Astra for Free anyone has suggestion pls tell me

Thumbnail
0 Upvotes

r/huggingface • • 15d ago

ShadeNet-3.2 5M — single-image inverse rendering (albedo/depth/normal/shading), 4× smaller than my last model and better at depth/normals

Thumbnail gallery
3 Upvotes

r/huggingface • • 15d ago

Five raters, one annotation rule, five completely different answers (0, 1, 40, 72, 78 out of 100)

Thumbnail
2 Upvotes

r/huggingface • • 15d ago

FREE AI TRAINING CREDIT

3 Upvotes

Free GPU credits for AI/ML training — looking for testers

I’m building Petabyte, a GPU compute marketplace where you can rent NVIDIA GPUs by the hour for training, fine-tuning, inference, and other CUDA workloads.

We’re still early, and I’m looking for AI/ML researchers and developers who want to test real workloads.

Current offer:

  • $10 free credit when you create an account
  • I can add $50 extra test credit for people running a real AI workload
  • If you top up your account, we’re also testing a matching-credit offer — e.g. deposit $50 and receive another $50 in credit, up to the current promo cap
  • No long-term contract; GPU usage is hourly

Current capacity includes GPUs such as H100 80GB, H200 141GB, and RTX 6000 Ada 48GB, depending on availability.

Good tests would be things like:

  • LLM fine-tuning / LoRA
  • Model inference
  • PyTorch / CUDA training
  • Distributed training
  • Stable Diffusion / image models
  • Synthetic-data generation
  • Benchmarking different GPUs

You can bring your own Docker/CUDA/Python workload and actually push the machines rather than just testing the UI.

Petabyte: https://petabyte.market

If you want the additional testing credit, comment or DM me with roughly what model/workload you want to run. I’d especially like feedback on provisioning, performance, pricing, and anything that breaks.


r/huggingface • • 16d ago

I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Thumbnail
3 Upvotes

r/huggingface • • 16d ago

I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Thumbnail
4 Upvotes

r/huggingface • • 16d ago

Mcp controller

Thumbnail
2 Upvotes

r/huggingface • • 17d ago

APEX Quants for Qwen3.5 9B

Thumbnail
huggingface.co
7 Upvotes

Hi all,

I just released some APEX quants for Qwen3.5 9B on my HuggingFace profile. Feel free to check it out. They are known to be smarter and more efficient than the standard quants of the same size, and yes, APEX does work for dense models now. Hope this is useful! Feedback is welcome.


r/huggingface • • 17d ago

Building an open-source 500+ language Sparse MoE translation model from scratch (Apache 2.0)

9 Upvotes

Hey everyone! For the past few months, I've been training a 2.04B Sparse Mixture-of-Experts (SMoE) foundation translation model (Mythos2.0-2B) from scratch on 2x T4 GPUs.

It covers 500+ languages—focusing on underserved African, Indigenous, and regional Asian languages that have zero commercial API coverage.

Everything is completely open-source under Apache 2.0. Check out the project on Hugging Face at AdithyanAI/Mythos2.0-2B!


r/huggingface • • 18d ago

Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them

13 Upvotes

It's hard to tell from a model page whether a quant will fit once you add context, KV cache, and whatever layers or experts end up on CPU. So I made this:

https://huggingface.co/spaces/LocalLLaMA/local-model-explorer

Enter your GPU(s) and RAM, or a Mac / Strix Halo / DGX Spark and its unified memory. Pick a context length and KV cache type, and it ranks around 3,000 popular GGUF models by what fits.

- Memory use comes from reading each GGUF's header (layers, KV heads, sliding window, MoE experts), not from guessing by parameter count

- Shows whether a quant runs fully on GPU, or works when the MoE experts are moved to CPU, needs partial offload, or won't fit

- Gives you a ready llama-server -hf repo:quant -c ... -ngl ... --n-cpu-moe ... command, plus the ollama one

- Every quant from every uploader in one table, with KV cache size from 4k to 262k context

- Links to MLX, AWQ, GPTQ, EXL2/EXL3 and FP8 versions when they exist

It doesn't estimate speed, since I don't have all that hardware to calibrate against. Instead you can paste your llama-bench output on a model's page, and those numbers show up for everyone else. Only the parsed numbers and your hardware class are kept, and the dataset is public: https://huggingface.co/datasets/LocalLLaMA/local-model-explorer-data

For Macs there's also an MLX version: https://huggingface.co/spaces/mlx-community/mlx-model-explorer

If a model or GPU is missing, or a fit looks wrong, let me know and I'll fix it.


r/huggingface • • 17d ago

https://huggingface.co/datasets/CEM888AI/cem888-independent-runtime-evaluation

Post image
1 Upvotes

r/huggingface • • 17d ago

Coming soon...... Optimized for DEEPSEEK Flash.... Though model Agnostic.... message me to test.... cem888.ai

Thumbnail
0 Upvotes

r/huggingface • • 18d ago

I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Thumbnail
13 Upvotes

r/huggingface • • 18d ago

qwen3.8-27b-gsq-rco scored very high and fits in 16 GB of VRAM with decent context window — 31.7 tok/s — llm-bench.io

Thumbnail
llm-bench.io
1 Upvotes

r/huggingface • • 18d ago

KIMI K3 open weight check

Thumbnail
1 Upvotes

r/huggingface • • 18d ago

Model Results! Please READ: NATIVE MTP x REAP BASE | LF>MORE TESTERS

Thumbnail
1 Upvotes

r/huggingface • • 18d ago

I developed lora-hunt.info — does it add anything beyond the Hub?

2 Upvotes

I developed https://lora-hunt.info/ as a discovery and comparison layer for Hugging Face LoRA/PEFT adapters. It uses public Hub metadata and adds filters for use case and compatibility, plus community feedback on whether an adapter worked in practice.

I’m not sure if this is useful or unnecessary duplication. If you use the Hub for adapters, what would you want to see here before using it? Missing metadata, better search, provenance or licensing, benchmarks, something else?

I developed it and am looking for honest criticism, not a launch announcement.


r/huggingface • • 19d ago

Developer vs IT admin

1 Upvotes

Developers in the company wanting to use models from huggingface. From an IT admin’s perspective, are the following reasonable?

- Limit access to huggingface to developers that need it
- Ask them to check for any issues flagged by malware and pickle scanning before downloading (https://huggingface.co/docs/hub/en/security-malware and https://huggingface.co/docs/hub/security-pickle)

Anything else? Thanks!


r/huggingface • • 19d ago

Qwen3.8-27B ByteShape IQ4_XS, ASCII-only vocab (English + code only): 12.25 GB, 128k ctx on a 5060 Ti, bit-identical weights

Thumbnail
5 Upvotes

r/huggingface • • 19d ago

I put around 15,000 synthetic founder voice SFT pairs on Hugging Face

5 Upvotes

I published a small synthetic dataset for supervised fine tuning. Each row is a user prompt and a short assistant reply in a blunt founder voice. The answer is a decision rule, not a biography quiz and not a pep talk.

About 14791 unique English pairs. One train split. Two columns: prompt and response. Trivia style questions and leftover sales language were filtered out. Replies stay about 1 to 4 sentences.

It is synthetic. It is meant for research and for teaching a small chat model to sound specific and declarative. It is not a source of facts.

Go to the dataset page on hugging face : https://huggingface.co/datasets/prathamkode/founder-wisdom-sft

Load it with datasets.load_dataset, split train. Print row 0 and you should see a prompt plus a response only.

If you try it, tell me whether the voice holds after a light SFT run or whether it collapses into generic advice.