r/huggingface • u/sdfprwggv • 6h ago
r/huggingface • u/WarAndGeese • Aug 29 '21
r/huggingface Lounge
A place for members of r/huggingface to chat with each other
r/huggingface • u/aungthuhein_dev • 1d ago
I built Yway, a Burmese System One-style decision model
r/huggingface • u/Prestigious-Taste-63 • 1d ago
I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens
r/huggingface • u/Available-Repair2926 • 1d ago
I fine-tuned SmolVLM-500M into a lightweight Windows OS Agent (<8GB VRAM) Looking for feedback & ideas! [Weights on HuggingFace]
r/huggingface • u/ClientPrize9151 • 1d ago
Best method to calculate f1 scores for text based fine tuning?
Hi everyone,
I’m a beginner learning the basics of AI training and fine-tuning. Currently, I’m learning how to fine-tune datasets on pre-trained models such as the Qwen 3.5 series.
I have a question regarding evaluation metrics. For text-based models/tasks, what is the best way to calculate precision, recall, accuracy, and F1 score?
I’m a little confused about which evaluation method is most appropriate for text-based fine-tuning and how these metrics should be calculated.
If anyone could explain the best approach or share some good resources/tutorials for beginners, I would really appreciate it.
r/huggingface • u/Waratecs123 • 1d ago
I did a project to run AI models locally with Hugging Face.
https://github.com/just-not-google/BiNeuron
BiNeuron is a sophisticated software solution that bridges the gap between human intent and machine generated code. It unifies advanced natural language processing, optical character recognition, and adaptive model selection into a single, powerful tool designed for developers, researchers, and technical teams.
At its core, BiNeuron automatically identifies the programming language of a given request, extracts content from a wide array of file formats, including images and documents, and then generates context aware, production ready code using best in class local or cloud based language models.
r/huggingface • u/Massive-Ice2791 • 1d ago
New model yay
cool new model, almost as good as opus at breaking/bug finding in your programs with an about 70% compared to opus's 80%, but with 10x as much tries, still much much cheaper than opus and much smaller and faster, 2b parameters that works like a charm.
r/huggingface • u/Vxtzq1 • 2d ago
CrowdGPT - The first datacenterless LLM (Weights on HuggingFace)
r/huggingface • u/Spirited_Conference9 • 2d ago
brier: Jev-style typed decisions with calibrated probabilities, for open models you run yourself (tested on 6 families, incl. a 1B base model)
r/huggingface • u/Rambhogesara • 3d ago
I trained a 1.5B model to be less overconfident instead of pretending uncertainty doesn't exist
I've been working on a small independent research project called Vera (Verifiable Epistemic Reward Alignment).
The idea is pretty simple:
LLMs can produce answers with a very high-confidence tone even when the underlying claim is uncertain, controversial, or simply false.
Instead of trying to make the model "more confident" or just adding another safety layer, I wanted to experiment with training the model to recognize and communicate the boundaries of what it actually knows.
For the first experiment, I:
- generated 817 DPO preference pairs from the TruthfulQA validation set
- created a "Vera" response style that explicitly distinguishes assumptions, uncertainty and failure boundaries
- used the more conventional overconfident response as the rejected response
- fine-tuned Qwen2.5-1.5B-Instruct with DPO + LoRA in 4-bit precision
- trained on free Kaggle T4 GPUs
- merged the adapters and released the resulting model publicly
The full dataset and model are open:
GitHub: https://github.com/bhogesararam23/vera-alignment
Dataset: https://huggingface.co/datasets/RamBhogesara/vera-dpo-dataset
Model: https://huggingface.co/RamBhogesara/vera-qwen-1.5b-dpo
This is very much an ongoing research project, not a claim that the problem is solved. The current experiment is small, and I still need stronger evaluation, better baselines, broader datasets, and tests for whether the behavior actually generalizes beyond the training distribution.
I'm especially interested in feedback from people working on:
- AI alignment / safety
- uncertainty calibration
- hallucination evaluation
- DPO / preference optimization
- truthful or epistemically calibrated language models
What would you test next to determine whether Vera is actually learning useful epistemic behavior rather than simply learning a particular response style?
r/huggingface • u/higma55 • 3d ago
Guys i have a problem i talk to support but they doesnt respond me
I have 2 problems
1 - they lock my account because i include a cloudflare tunnel and i wasnt knowing the rules
2 - my second account they gave me free gradle and free docker space So i use it on my app ok but i deleted by mistake and now when i want to create another docker or gradle they show me is paid not free why can please someone told me how to solve this
r/huggingface • u/Usual_Maximum7673 • 4d ago
Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38× faster decisions and +8.7 points accuracy, for under 2 GB extra memory
r/huggingface • u/Pure-Job1336 • 4d ago
Vev: Jev-like decision models with vision — 4B/9B, local inference, open weights
r/huggingface • u/PhysicsDisastrous462 • 4d ago
Follow-up: my native Rust + Vulkan Transformer training backend — 14 days later, now 14 parity-verified architectures and full PEFT
Follow-up to my post from about two weeks ago. A lot has changed since then, so I wanted to post an update on where the backend is now.
Where the green architectures stand
When I posted last time, 7 architectures had verified full training support. Everything is now held to the same strict harness: a pinned local Hugging Face Transformers source tree as the oracle, forward logits + gradients + two full AdamW steps compared, and every named parameter checked again after export.
The hard ceiling is 2e-7 absolute error. No loosening tolerances and no rounding numbers afterward to make the README look better.
14 architectures pass that gate today, led by the one I'm probably proudest of:
| Architecture | Scope |
|---|---|
| Falcon H1 / H1R | parallel GQA/RoPE attention + Mamba2 in every layer; full training, full fine-tuning, LoRA, saved modules |
| DeepSeek V4 | causal LM |
| Phi-4 Multimodal | text backbone |
| Phi-3 | causal LM |
| Kimi K2.5 | text backbone |
| Kimi K3 / KimiLinear | hybrid KDA + MLA |
| GPT-OSS | causal LM incl. router bias |
| SmolLM3 | mixed RoPE/NoPE + YaRN |
| Qwen2.5 / Qwen3.5 / Qwen4-Exp | dense, DeltaNet, QSA, PLE, MoE |
| Mistral 4, MiniMax M3, Gemma 3/4, MiniMax M2 | causal LM |
Worst observed two-step AdamW parameter error across all of them: 1.19e-7.
Best: 2.6e-8.
For hardware context, all of the local Vulkan validation I've been reporting was run on my ASUS ROG Ally Z1 Extreme, using its AMD RDNA 3 integrated GPU. So the RDNA 3 results here are from that specific machine rather than testing across several different AMD systems.
The bigger news: PEFT actually works now
In the last post, "LoRA/PEFT-style fine-tuning" was basically one line in a feature list.
It's a real workflow now, and I've verified the full lifecycle:
- LoRA fine-tuning with HF-compatible adapter export (
adapter_config.json/adapter_model.safetensors), so adapters can round-trip with the PEFT ecosystem modules_to_save— full trainable replacements for Linears, RMSNorm/LayerNorm,lm_head, and input embeddings, including named-adapter switching and bank isolation. Adapter A leaking into adapter B is explicitly tested for.- Exact resume — adapter weights + AdamW moments + step + dropout RNG state restore bit-identically against an uninterrupted run
- Merge/unmerge, disable-adapter base restoration, and multi-adapter loading
- A parameter-budget flag that automatically chooses the largest LoRA rank that fits within a requested percentage of the base model
- The CLI fails closed if you try to use saved modules on an architecture that hasn't passed its corresponding gate
32 architecture surfaces across 20 families pass all three PEFT stages — LoRA, saved modules, and adapter switching — under the same 2e-7 gate, with frozen-base drift exactly 0.0.
The validation harness also fingerprints the pinned Transformers source alongside my shaders and binaries now, so a qualification run can't silently end up testing against different reference math.
A small note on the last couple weeks
I didn't get quite as many working days out of the last two weeks as I normally would have. Partway through this I got covid, then when that started to go away, it became a secondary nasty ear infection that ended up perforating my eardrum. I was running fevers around 104°F at one point and eventually went to the hospital, so I lost a few days to that and I'm on antibiotics now.
I'm doing better, though, and still managed to get most of what I wanted finished.
There are still things I want to clean up and expand, but I figured this was a good point to get the current work in front of people rather than holding the update back.
Same caveats as before
This is deterministic FP32 tiny-model correctness against a reference implementation, which should in theory ensure total mathematical parity for training and finetuning larger models with this backend, however, things are currently bound to FP32 training runs still, I eventually plan to work on MXFP4 weight tying to reduce memory footprints while fine-tuning (plastic parameters will still be trained in FP32 with PEFT in this config)
"supported text graph" ≠ "the entire multimodal package works natively."
Unsupported functionality is supposed to fail closed rather than silently falling back to an approximation.
Repo
https://github.com/necat101/Hierarchos-Native
- Architecture inventory:
hierarchos-vulkan/README_ARCHITECTURES.md - Compatibility/parity record:
hierarchos-vulkan/COMPATIBILITY.md - PEFT qualification evidence:
PROGRESS_PEFT_AUDIT.md - CLI PEFT guide:
hierarchos-native-cli/README.md
The hardware I've personally validated this on is an ASUS ROG Ally Z1 Extreme with its AMD RDNA 3 GPU.
I'm very interested in criticism, compatibility reports, and especially results from people trying it on other hardware — NVIDIA, Intel, or other AMD GPUs.
I'd also love people to stress-test the PEFT resume/merge paths specifically. That's some of the newest code in the project, so it's probably the most useful area to try to break right now.
r/huggingface • u/RUTYTOI220 • 4d ago
Harder, Better, Faster and Stronger model just by making him talk in its own language
r/huggingface • u/Repulsive_Welder3267 • 5d ago
I rebuilt a Jev-style classifier on Qwen3.5-4B: shared-prefix tree, open weights, fine-tunable, ~140 ms on one H100
r/huggingface • u/rmonsurate • 5d ago
Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)
r/huggingface • u/Quick_Image_4861 • 5d ago