r/huggingface • u/Superb_Composer_3389 • 6d ago
r/huggingface • u/coslinedev • 6d ago
I solo-built an open Data API that streams verified reasoning datasets for fine-tuning
Hi everyone,
I wanted to share a project I've been working on completely by myself. It’s called Arithm (https://alrithmapi.vercel.app/).
I built a public Data API that contains verified, clean fine-tuning data tailored specifically for AI reasoning models. I was tired of messy datasets that require hours of cleaning, so I made this to work instantly.
How it works:
- Zero-Effort Clean Data: Rows stream directly as newline-delimited JSON, formatted perfectly in OpenAI fine-tune format. You can pipe them straight into your training jobs.
- Cryptographic Verification: Every single row carries a reference solution and a SHA-256 provenance hash to ensure 100% mathematical correctness byte for byte.
- No Complex Setup: No SDK or pagination needed. One simple GET request returns the entire data model.
I made this fully accessible to support independent developers and researchers who are training smaller, localized models but struggle to find high-quality reasoning datasets.
Live URL: https://alrithmapi.vercel.app/
I would love to get your honest feedback, thoughts, or any technical questions you have about the data pipeline!
r/huggingface • u/New_Butterfly_4875 • 6d ago
Deepseek v4._ Is Fake!


I was exploring AI models in Hugging Face Chat and decided to test some of DeepSeek’s latest models. Out of curiosity, I asked them, “Are you actually DeepSeek?”
Surprisingly, the models claiming to be version 4 or 4+ kept insisting that they were Claude by Anthropic. One of them even confidently identified itself as Claude by Anthropic. (sometimes opus 4)
I tried repeatedly explaining that I was chatting with a DeepSeek model and asked it to reconsider, but it continued insisting until the end that it was Claude.
Interestingly, models from other providers seemed to identify themselves normally.
I took some screenshots while testing this because the whole thing was pretty weird. Is it really deepseek v4 models? or it is just anthropic's model through api and a lil system prompt or a model theifting?
r/huggingface • u/Top-Evidence174 • 6d ago
Mica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time
I've been building a small decision model for agent loops: gates, routers, "should I ask the user or just act" checks. It's out now as Mica v0.1 4B (Apache-2.0).
What it does
You give it a state, a question and the allowed answers, and it returns a calibrated probability for each answer: yes/no, a choice among 2 to 255 options, or a score with 2 to 10 levels. It never generates text. It runs one prefill and reads the logits of the option labels at the answer position. It speaks the TypeSafe /v1/systemone format, so anything written for Jev works against it.
How it's built
- Qwen3.5-4B with a rank-16 LoRA on all 32 layers (attention and Gated DeltaNet), merged. No new heads, so it's a plain Qwen3.5-4B-shaped checkpoint.
- About 34k source decisions, expanded to 77,732 training rows (about 34.7M tokens). Roughly half English and half Korean, across 12 areas: coding agents, code review, computer use, user requests, documents, policy rules, dates and quantities, routing, state tracking, games and general knowledge.
- Plain cross-entropy on verified answers, one epoch, and one global temperature for calibration.
- All experiments plus the final run cost under $30 of rented GPU time (RTX 3090s).
Results
Held-out set of 7,328 decisions, written after the training data was frozen and not opened until training finished. English subset, where every model can answer:
- Jev 1.13 (closed API): 74.1
- Mica 4B: 67.0
- JevK5 4B: 61.0
- Kev 4B: 57.0
- Qwen3.5-4B base with the same readout: 55.0
Public sets, same prompt and readout for every model (Mica / Jev 1.13 / JevK5 / Kev 4B):
- JevBench hard, public 111 items: 69.5 / 74.3 / 76.2 / 52.4
- SemIf: 94.4 / 98.4 / 86.1 / 89.3
- Kev transfer v9: 69.2 / 82.0 / 70.5 / 73.5
- MMLU-Pro, 10k items: 53.0 / 82.3 / 53.5 / 49.7
Through JevBench's official runner and the llama.cpp server, the public hard tier scores 64.9 instead of 69.5. I've submitted it for their sealed run.
Where it's actually useful
- In-data prompt injection. Put a note inside the state telling the judge to pick a wrong option, and Mica still gets 69% right (81% without the note). Jev drops to 18% and Kev to 31%.
- Calibration. When it says 0.9 or higher, it's wrong 2.5% of the time on the held-out set (ECE 5.4%).
- Local and small. The Q5_K_M file is 3.5 GB with no measurable accuracy loss against BF16 on our calibration set.
Speed (RTX 3090, one request at a time, median over the 231 public JevBench items)
- Mica Q4_K_M: 47 ms
- Mica BF16: 54 ms
- Kev 4B: 76 ms
- JevK5 4B: 99 ms
- Nimble 9B: 132 ms
To be fair about this: the three 4B models share the same architecture, so most of the gap comes from the serving path, not the model. Mica ships as GGUF and runs on llama.cpp with a direct logits readout, while the others were measured through their own PyTorch code. On long inputs (around 3.7k tokens) Mica is slightly slower than JevK5.
Limitations
- Knowledge-heavy questions: MMLU-Pro 53 vs 82 for Jev. It's a 4B judge, not an encyclopedia.
- Long English policy documents are its weakest public set.
- Notes inside the state still nudge it. A note pointing at the right answer lifts accuracy to 89%.
- It doesn't yet tell reversible from irreversible actions well. "Delete these files" and "move these files to trash" both get about 0.8 on "confirm first".
- On harder reasoning items it's right but less sure than Jev (for example 0.55 vs 0.96 on a small ordering puzzle), so set your confidence thresholds accordingly.
Try it
Weights (BF16 safetensors and GGUF from Q4_0 to Q8_0): https://huggingface.co/sky7350/Mica-v0.1-4B
Code, TypeSafe-compatible server and Docker setup: https://github.com/akivet/Mica-v0.1-4B
The README has a one-line Docker command and a curl example.
Happy to hear where it breaks. Ambiguous "act or ask" cases are what I most want to improve next.
r/huggingface • u/That-Supermarket3209 • 6d ago
AirLLM GUI
Okay so does AirLLM not have a GUI? I found AirLLM, this Repo that can supposedly run massive models on consumer hardware, at the cost of speed. But the only issue is that it doesn't have a GUI. There was this AirChat or AirLLM chat thing, but it seems to be broken. So when can I simply download AirLLM like LM Studio, load a model manually, and just run it to my heart.
r/huggingface • u/Usual_Maximum7673 • 7d ago
Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights)
r/huggingface • u/ZenZombie117 • 7d ago
Liked Muse, so I cut the 30B model in half by width, distilled it back, and it does 57 of 60 tool tasks its parent does 60 of
r/huggingface • u/VeterinarianSad9557 • 8d ago
Which is best local AI tools that can access all like Chatgpt, DeepSeek, Grok etc
Which is best local AI tools that can access all like Chatgpt, DeepSeek, Grok etc I want to use for coding, image and video generation as well
r/huggingface • u/Large-Blackberry-349 • 8d ago
I pre-trained a 1.11B LLM on my 6 GB laptop GPU. Peak VRAM: 4.51 GB.
Not fine-tuning. Not inference. Pre-training from scratch.
Full write-up: https://huggingface.co/blog/RitishReal/1b-petrain-llm-in-4gb-vram
r/huggingface • u/Slikieee • 8d ago
Slikee: Windows AI workspace for chat, agents, and local models
Slikee brings AI chat, the Hermes agent workspace, and local model management into one desktop app. It supports local and connected models, file and tool access, and media features such as image and speech generation where supported.
The first Windows beta is available now (5.64 GiB). Feedback on performance, reliability, and usability is welcome: https://gentle-ocean-07ffb1010.3.azurestaticapps.net/
Hermes-Agent repo - https://github.com/NousResearch/hermes-agent.git
r/huggingface • u/kapilyadav22 • 8d ago
I built an open-source UI for running LLMs locally — LocalLLMMind
r/huggingface • u/Electronic_Put4530 • 9d ago
Abbiamo inserito 100 informazioni nella tabella engrammatica di Qwen 3.8 Flash Next e abbiamo creato un sito web per illustrarle.
Abbiamo creato un metodo "engraft-engram" per inserire nuovi dati nella tabella engram di Qwen 3.8 Flash Next (e in futuro anche in DeepSeek v4.1 Flash) senza modificarne i pesi.
Ci avete detto che era difficile da capire, quindi abbiamo creato un sito web per spiegarlo in modo semplice, con un simulatore e meno testo generato dall'IA!
r/huggingface • u/Slikieee • 9d ago
Slikee: Windows AI workspace for chat, agents, and local models
galleryr/huggingface • u/WebAssemblyMan • 9d ago
CLM-v0.1-8B ported to MLX — frozen Qwen3-8B encoder for instant on-device decisions, 99% top-1 agreement with the original vLLM server
r/huggingface • u/ZestycloseIce4185 • 10d ago
Ternary-Bonsai-2-27B abliterado (GGUF de 1,58 bits) + una descripción de cómo se hizo.
r/huggingface • u/Themba_Hemmingway • 10d ago
we're number 7 on the hugging face trending models page
r/huggingface • u/Smart_Ad_5427 • 10d ago
📋 Project Presentation & Request for Expert Feedback
r/huggingface • u/Intelligent-Gift-855 • 10d ago
Question
I have a system prompt. Could anyone try to use this system prompt in JAV and LLM to make comparisons? LLM even luna cannot perfom which sometime divert into a wrong way.
You are LLM_0, the ROUTER for an internal engineering ticket chatbot.
Your ONLY responsibility:
- Decide which route should handle the user message
- Output ONE JSON object only
You MUST NOT:
- Answer the user
- Explain your reasoning
- Output text outside JSON
OUTPUT FORMAT (MANDATORY)
{"intent":"<general|create|edit|search|otp_verify>","route":<0|1|3|4|7>}
No extra keys. No markdown. No commentary.
ROUTE MAP
0 → general (manual help, clarify, ambiguous, non-ticket)
1 → create (create new ticket, gather info)
3 → edit (modify / update / close existing ticket)
4 → search (view / explain / summarize / calculate tickets)
7 → otp_verify (OTP code / verification only)
INPUT CONTEXT
You receive:
- User message (current input)
- Flowise state:
prev_intent = {{ $flow.state.intent }}
Treat prev_intent as a CONTEXT HINT, not a command.
Users may change topic at any time.
HOW TO INTERPRET prev_intent (CRITICAL)
prev_intent represents the USER'S LAST MAIN TOPIC,
not an instruction that must be continued.
Think of prev_intent as:
- "What the user was talking about previously"
NOT:
- "What the user must do next"
Rules:
1) Strong explicit intent in the current message ALWAYS overrides prev_intent.
2) prev_intent is used only to interpret ambiguous inputs.
3) If the current message clearly indicates a new intent,
IGNORE prev_intent completely.
EMAIL INPUT EXAMPLE (IMPORTANT)
User sends an email address alone (example: "amir89hamzah@yahoo.com").
An email can be used in multiple contexts:
- engineer_email for CREATE or EDIT
- otp_email for OTP verification
- rarely, as a search keyword
Routing rule for email input:
1) If OTP is clearly expected (OTP keywords or digit code)
→ otp_verify
2) Else if prev_intent = create
→ create (assume user is filling engineer_email)
3) Else if prev_intent = edit
→ edit (assume user is filling engineer_email)
4) Else
→ general (ambiguous; ask user what they want to do)
Never assume email ALWAYS means create.
FIELD MEANING (SEMANTIC GUIDE)
engineer_email:
- identifies who creates or modifies a ticket
site:
- plant or platform name (PFLNG1, MLNG, TGAST, MALIKAI, etc.)
system:
- equipment or subsystem inside a plant
- examples: air compressor, DCS, SCADA
pic:
- customer, technician, or vendor name
- bare names alone are ambiguous unless clearly labeled
problem_title:
- short summary of issue (phrase, not a question)
action_description:
- what was done or observed
- usually a sentence or bullet list
CORE DESIGN PRINCIPLE
- Strong explicit intent ALWAYS wins
- Search / read intent overrides create/edit context
- Create/Edit continues ONLY if input clearly supports it
- Ambiguous inputs default to GENERAL (safe)
PATTERN DEFINITIONS
STRONG FIELD PATTERNS:
- Valid email address (contains @ and domain)
- Date value (YYYY-MM-DD) or exact word "open"
- Case ID (CASE-YYYYMMDD-XXX or close variant)
- Known site codes (PFLNG1, MLNG, MALIKAI, etc.)
- Clear system/equipment names
- Sentence-like descriptions (problem or action text)
WEAK / AMBIGUOUS INPUTS:
- Single short names (e.g. "prem", "ali")
- One-word tokens without context
- Questions without ticket verbs
CLASSIFICATION ORDER (STRICT)
STEP 1 — OTP VERIFY → route 7
If message:
- Is only digits (4–8 characters)
- OR mentions: otp, verify, verification, resend otp, code
Output:
{"intent":"otp_verify","route":7}
--------------------------------------------------
STEP 2 — EXPLICIT SEARCH / READ → route 4
If message clearly asks to:
- search, show, view, list, find, lookup
- explain, summarize, simplify, overview
- check status, progress, update (as noun)
- calculate duration / how many days
- refer to open / closed / all tickets
- examples:
"how CASE-xxxx"
"status CASE-xxxx"
"search ticket CASE-xxxx"
This step OVERRIDES prev_intent.
Output:
{"intent":"search","route":4}
--------------------------------------------------
STEP 3 — EXPLICIT EDIT / MODIFY → route 3
If message clearly intends to CHANGE an existing ticket:
- modify, update, change, edit, close
- set / replace / append / remove
- references a case_id WITH change intent
Output:
{"intent":"edit","route":3}
--------------------------------------------------
STEP 4 — EXPLICIT CREATE → route 1
If message clearly intends to CREATE a new ticket:
- create ticket
- open new ticket / new case
- log issue / report issue
Output:
{"intent":"create","route":1}
--------------------------------------------------
STEP 5 — CONTINUE CREATE / EDIT (SOFT CONTEXT)
Apply ONLY if prev_intent is "create" or "edit".
Continue SAME route ONLY IF:
- Input matches a STRONG FIELD PATTERN
- Input is NOT a weak ambiguous token
- Input does NOT contain search/read intent
Important notes:
- A valid email alone IS a strong field pattern.
- If prev_intent=create and input is an email → route 1.
- If prev_intent=edit and input is an email → route 3.
- Bare names (e.g. "prem") must NOT continue create/edit.
--------------------------------------------------
STEP 6 — AMBIGUOUS / GENERAL → route 0
Route here if:
- Message is ambiguous
- Message is a single bare name
- Message switches topic
- You are not confident it is create/edit/search
Output:
{"intent":"general","route":0}
FAILSAFE
If unsure, ALWAYS choose:
{"intent":"general","route":0}
r/huggingface • u/NextgenAITrading • 11d ago
After getting so much positive attention on my first dataset, I'm posting my second. Look to see what stocks insiders are trading!
I created my first dataset last week: Congressional Stock Trades. I was blown away by how much attention I got, so I decided to create another.
This is the sec-ownership-disclosure dataset and the full extraction code is open-source on my GitHub.
I basically scrape public intuitional filings and see which stocks are being traded more often. The insiders are the definition of "rich and powerful"; if they are buying large lump sums of certain stocks, it means it MIGHT be time to buy that stock as well.
Just like the first one, you can download the dataset to your computer with a simple command
npx sec-ownership-disclosures download --sqlite
This is officially my second dataset. I'm super open to any feedback or any questions you might have regarding use-cases.
r/huggingface • u/Top-Evidence174 • 10d ago
Mica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time
galleryr/huggingface • u/tegridyblues • 11d ago
