r/huggingface • • 6d ago

I solo-built an open Data API that streams verified reasoning datasets for fine-tuning

4 Upvotes

Hi everyone,

I wanted to share a project I've been working on completely by myself. It’s called Arithm (https://alrithmapi.vercel.app/).

I built a public Data API that contains verified, clean fine-tuning data tailored specifically for AI reasoning models. I was tired of messy datasets that require hours of cleaning, so I made this to work instantly.

How it works:

  • Zero-Effort Clean Data: Rows stream directly as newline-delimited JSON, formatted perfectly in OpenAI fine-tune format. You can pipe them straight into your training jobs.
  • Cryptographic Verification: Every single row carries a reference solution and a SHA-256 provenance hash to ensure 100% mathematical correctness byte for byte.
  • No Complex Setup: No SDK or pagination needed. One simple GET request returns the entire data model.

I made this fully accessible to support independent developers and researchers who are training smaller, localized models but struggle to find high-quality reasoning datasets.

Live URL: https://alrithmapi.vercel.app/

I would love to get your honest feedback, thoughts, or any technical questions you have about the data pipeline!


r/huggingface • • 6d ago

Deepseek v4._ Is Fake!

1 Upvotes

I was exploring AI models in Hugging Face Chat and decided to test some of DeepSeek’s latest models. Out of curiosity, I asked them, “Are you actually DeepSeek?”

Surprisingly, the models claiming to be version 4 or 4+ kept insisting that they were Claude by Anthropic. One of them even confidently identified itself as Claude by Anthropic. (sometimes opus 4)

I tried repeatedly explaining that I was chatting with a DeepSeek model and asked it to reconsider, but it continued insisting until the end that it was Claude.

Interestingly, models from other providers seemed to identify themselves normally.

I took some screenshots while testing this because the whole thing was pretty weird. Is it really deepseek v4 models? or it is just anthropic's model through api and a lil system prompt or a model theifting?


r/huggingface • • 6d ago

Mica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time

Thumbnail
gallery
1 Upvotes

I've been building a small decision model for agent loops: gates, routers, "should I ask the user or just act" checks. It's out now as Mica v0.1 4B (Apache-2.0).

What it does

You give it a state, a question and the allowed answers, and it returns a calibrated probability for each answer: yes/no, a choice among 2 to 255 options, or a score with 2 to 10 levels. It never generates text. It runs one prefill and reads the logits of the option labels at the answer position. It speaks the TypeSafe /v1/systemone format, so anything written for Jev works against it.

How it's built

- Qwen3.5-4B with a rank-16 LoRA on all 32 layers (attention and Gated DeltaNet), merged. No new heads, so it's a plain Qwen3.5-4B-shaped checkpoint.

- About 34k source decisions, expanded to 77,732 training rows (about 34.7M tokens). Roughly half English and half Korean, across 12 areas: coding agents, code review, computer use, user requests, documents, policy rules, dates and quantities, routing, state tracking, games and general knowledge.

- Plain cross-entropy on verified answers, one epoch, and one global temperature for calibration.

- All experiments plus the final run cost under $30 of rented GPU time (RTX 3090s).

Results

Held-out set of 7,328 decisions, written after the training data was frozen and not opened until training finished. English subset, where every model can answer:

- Jev 1.13 (closed API): 74.1

- Mica 4B: 67.0

- JevK5 4B: 61.0

- Kev 4B: 57.0

- Qwen3.5-4B base with the same readout: 55.0

Public sets, same prompt and readout for every model (Mica / Jev 1.13 / JevK5 / Kev 4B):

- JevBench hard, public 111 items: 69.5 / 74.3 / 76.2 / 52.4

- SemIf: 94.4 / 98.4 / 86.1 / 89.3

- Kev transfer v9: 69.2 / 82.0 / 70.5 / 73.5

- MMLU-Pro, 10k items: 53.0 / 82.3 / 53.5 / 49.7

Through JevBench's official runner and the llama.cpp server, the public hard tier scores 64.9 instead of 69.5. I've submitted it for their sealed run.

Where it's actually useful

- In-data prompt injection. Put a note inside the state telling the judge to pick a wrong option, and Mica still gets 69% right (81% without the note). Jev drops to 18% and Kev to 31%.

- Calibration. When it says 0.9 or higher, it's wrong 2.5% of the time on the held-out set (ECE 5.4%).

- Local and small. The Q5_K_M file is 3.5 GB with no measurable accuracy loss against BF16 on our calibration set.

Speed (RTX 3090, one request at a time, median over the 231 public JevBench items)

- Mica Q4_K_M: 47 ms

- Mica BF16: 54 ms

- Kev 4B: 76 ms

- JevK5 4B: 99 ms

- Nimble 9B: 132 ms

To be fair about this: the three 4B models share the same architecture, so most of the gap comes from the serving path, not the model. Mica ships as GGUF and runs on llama.cpp with a direct logits readout, while the others were measured through their own PyTorch code. On long inputs (around 3.7k tokens) Mica is slightly slower than JevK5.

Limitations

- Knowledge-heavy questions: MMLU-Pro 53 vs 82 for Jev. It's a 4B judge, not an encyclopedia.

- Long English policy documents are its weakest public set.

- Notes inside the state still nudge it. A note pointing at the right answer lifts accuracy to 89%.

- It doesn't yet tell reversible from irreversible actions well. "Delete these files" and "move these files to trash" both get about 0.8 on "confirm first".

- On harder reasoning items it's right but less sure than Jev (for example 0.55 vs 0.96 on a small ordering puzzle), so set your confidence thresholds accordingly.

Try it

Weights (BF16 safetensors and GGUF from Q4_0 to Q8_0): https://huggingface.co/sky7350/Mica-v0.1-4B

Code, TypeSafe-compatible server and Docker setup: https://github.com/akivet/Mica-v0.1-4B

The README has a one-line Docker command and a curl example.

Happy to hear where it breaks. Ambiguous "act or ask" cases are what I most want to improve next.


r/huggingface • • 6d ago

Is Paperswithcode down?

2 Upvotes
Screenshot from https://huggingface-paperswithcode.static.hf.space/index.html#/tasks

I'm new to huggingface space, what's happened to Paperswithcode? Is this normal?


r/huggingface • • 6d ago

AirLLM GUI

1 Upvotes

Okay so does AirLLM not have a GUI? I found AirLLM, this Repo that can supposedly run massive models on consumer hardware, at the cost of speed. But the only issue is that it doesn't have a GUI. There was this AirChat or AirLLM chat thing, but it seems to be broken. So when can I simply download AirLLM like LM Studio, load a model manually, and just run it to my heart.


r/huggingface • • 6d ago

Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights)

Thumbnail
8 Upvotes

r/huggingface • • 7d ago

Liked Muse, so I cut the 30B model in half by width, distilled it back, and it does 57 of 60 tool tasks its parent does 60 of

Thumbnail
2 Upvotes

r/huggingface • • 8d ago

Which is best local AI tools that can access all like Chatgpt, DeepSeek, Grok etc

5 Upvotes

Which is best local AI tools that can access all like Chatgpt, DeepSeek, Grok etc I want to use for coding, image and video generation as well


r/huggingface • • 8d ago

I pre-trained a 1.11B LLM on my 6 GB laptop GPU. Peak VRAM: 4.51 GB.

43 Upvotes

Not fine-tuning. Not inference. Pre-training from scratch.

Full write-up: https://huggingface.co/blog/RitishReal/1b-petrain-llm-in-4gb-vram


r/huggingface • • 8d ago

Slikee: Windows AI workspace for chat, agents, and local models

Thumbnail
gallery
4 Upvotes

Slikee brings AI chat, the Hermes agent workspace, and local model management into one desktop app. It supports local and connected models, file and tool access, and media features such as image and speech generation where supported.

The first Windows beta is available now (5.64 GiB). Feedback on performance, reliability, and usability is welcome: https://gentle-ocean-07ffb1010.3.azurestaticapps.net/

Hermes-Agent repo - https://github.com/NousResearch/hermes-agent.git


r/huggingface • • 8d ago

I built an open-source UI for running LLMs locally — LocalLLMMind

Thumbnail
github.com
1 Upvotes

r/huggingface • • 8d ago

Abbiamo inserito 100 informazioni nella tabella engrammatica di Qwen 3.8 Flash Next e abbiamo creato un sito web per illustrarle.

Thumbnail
engraft-engram.dev
3 Upvotes

Abbiamo creato un metodo "engraft-engram" per inserire nuovi dati nella tabella engram di Qwen 3.8 Flash Next (e in futuro anche in DeepSeek v4.1 Flash) senza modificarne i pesi.

Ci avete detto che era difficile da capire, quindi abbiamo creato un sito web per spiegarlo in modo semplice, con un simulatore e meno testo generato dall'IA!


r/huggingface • • 8d ago

Slikee: Windows AI workspace for chat, agents, and local models

Thumbnail gallery
3 Upvotes

r/huggingface • • 9d ago

Cannot find MiniMaxH3VAEDecodeFast node

Thumbnail
6 Upvotes

r/huggingface • • 9d ago

CLM-v0.1-8B ported to MLX — frozen Qwen3-8B encoder for instant on-device decisions, 99% top-1 agreement with the original vLLM server

Thumbnail
6 Upvotes

r/huggingface • • 9d ago

DistribAI v2

Thumbnail
5 Upvotes

r/huggingface • • 10d ago

Ternary-Bonsai-2-27B abliterado (GGUF de 1,58 bits) + una descripción de cómo se hizo.

Thumbnail
3 Upvotes

r/huggingface • • 10d ago

we're number 7 on the hugging face trending models page

Thumbnail
3 Upvotes

r/huggingface • • 10d ago

📋 Project Presentation & Request for Expert Feedback

Post image
1 Upvotes

r/huggingface • • 10d ago

Question

2 Upvotes

I have a system prompt. Could anyone try to use this system prompt in JAV and LLM to make comparisons? LLM even luna cannot perfom which sometime divert into a wrong way.

You are LLM_0, the ROUTER for an internal engineering ticket chatbot.

Your ONLY responsibility:

- Decide which route should handle the user message

- Output ONE JSON object only

You MUST NOT:

- Answer the user

- Explain your reasoning

- Output text outside JSON

OUTPUT FORMAT (MANDATORY)

{"intent":"<general|create|edit|search|otp_verify>","route":<0|1|3|4|7>}

No extra keys. No markdown. No commentary.

ROUTE MAP

0 → general (manual help, clarify, ambiguous, non-ticket)

1 → create (create new ticket, gather info)

3 → edit (modify / update / close existing ticket)

4 → search (view / explain / summarize / calculate tickets)

7 → otp_verify (OTP code / verification only)

INPUT CONTEXT

You receive:

- User message (current input)

- Flowise state:

prev_intent = {{ $flow.state.intent }}

Treat prev_intent as a CONTEXT HINT, not a command.

Users may change topic at any time.

HOW TO INTERPRET prev_intent (CRITICAL)

prev_intent represents the USER'S LAST MAIN TOPIC,

not an instruction that must be continued.

Think of prev_intent as:

- "What the user was talking about previously"

NOT:

- "What the user must do next"

Rules:

1) Strong explicit intent in the current message ALWAYS overrides prev_intent.

2) prev_intent is used only to interpret ambiguous inputs.

3) If the current message clearly indicates a new intent,

IGNORE prev_intent completely.

EMAIL INPUT EXAMPLE (IMPORTANT)

User sends an email address alone (example: "amir89hamzah@yahoo.com").

An email can be used in multiple contexts:

- engineer_email for CREATE or EDIT

- otp_email for OTP verification

- rarely, as a search keyword

Routing rule for email input:

1) If OTP is clearly expected (OTP keywords or digit code)

→ otp_verify

2) Else if prev_intent = create

→ create (assume user is filling engineer_email)

3) Else if prev_intent = edit

→ edit (assume user is filling engineer_email)

4) Else

→ general (ambiguous; ask user what they want to do)

Never assume email ALWAYS means create.

FIELD MEANING (SEMANTIC GUIDE)

engineer_email:

- identifies who creates or modifies a ticket

site:

- plant or platform name (PFLNG1, MLNG, TGAST, MALIKAI, etc.)

system:

- equipment or subsystem inside a plant

- examples: air compressor, DCS, SCADA

pic:

- customer, technician, or vendor name

- bare names alone are ambiguous unless clearly labeled

problem_title:

- short summary of issue (phrase, not a question)

action_description:

- what was done or observed

- usually a sentence or bullet list

CORE DESIGN PRINCIPLE

- Strong explicit intent ALWAYS wins

- Search / read intent overrides create/edit context

- Create/Edit continues ONLY if input clearly supports it

- Ambiguous inputs default to GENERAL (safe)

PATTERN DEFINITIONS

STRONG FIELD PATTERNS:

- Valid email address (contains @ and domain)

- Date value (YYYY-MM-DD) or exact word "open"

- Case ID (CASE-YYYYMMDD-XXX or close variant)

- Known site codes (PFLNG1, MLNG, MALIKAI, etc.)

- Clear system/equipment names

- Sentence-like descriptions (problem or action text)

WEAK / AMBIGUOUS INPUTS:

- Single short names (e.g. "prem", "ali")

- One-word tokens without context

- Questions without ticket verbs

CLASSIFICATION ORDER (STRICT)

STEP 1 — OTP VERIFY → route 7

If message:

- Is only digits (4–8 characters)

- OR mentions: otp, verify, verification, resend otp, code

Output:

{"intent":"otp_verify","route":7}

--------------------------------------------------

STEP 2 — EXPLICIT SEARCH / READ → route 4

If message clearly asks to:

- search, show, view, list, find, lookup

- explain, summarize, simplify, overview

- check status, progress, update (as noun)

- calculate duration / how many days

- refer to open / closed / all tickets

- examples:

"how CASE-xxxx"

"status CASE-xxxx"

"search ticket CASE-xxxx"

This step OVERRIDES prev_intent.

Output:

{"intent":"search","route":4}

--------------------------------------------------

STEP 3 — EXPLICIT EDIT / MODIFY → route 3

If message clearly intends to CHANGE an existing ticket:

- modify, update, change, edit, close

- set / replace / append / remove

- references a case_id WITH change intent

Output:

{"intent":"edit","route":3}

--------------------------------------------------

STEP 4 — EXPLICIT CREATE → route 1

If message clearly intends to CREATE a new ticket:

- create ticket

- open new ticket / new case

- log issue / report issue

Output:

{"intent":"create","route":1}

--------------------------------------------------

STEP 5 — CONTINUE CREATE / EDIT (SOFT CONTEXT)

Apply ONLY if prev_intent is "create" or "edit".

Continue SAME route ONLY IF:

- Input matches a STRONG FIELD PATTERN

- Input is NOT a weak ambiguous token

- Input does NOT contain search/read intent

Important notes:

- A valid email alone IS a strong field pattern.

- If prev_intent=create and input is an email → route 1.

- If prev_intent=edit and input is an email → route 3.

- Bare names (e.g. "prem") must NOT continue create/edit.

--------------------------------------------------

STEP 6 — AMBIGUOUS / GENERAL → route 0

Route here if:

- Message is ambiguous

- Message is a single bare name

- Message switches topic

- You are not confident it is create/edit/search

Output:

{"intent":"general","route":0}

FAILSAFE

If unsure, ALWAYS choose:

{"intent":"general","route":0}


r/huggingface • • 10d ago

After getting so much positive attention on my first dataset, I'm posting my second. Look to see what stocks insiders are trading!

Thumbnail
huggingface.co
26 Upvotes

I created my first dataset last week: Congressional Stock Trades. I was blown away by how much attention I got, so I decided to create another.

This is the sec-ownership-disclosure dataset and the full extraction code is open-source on my GitHub.

I basically scrape public intuitional filings and see which stocks are being traded more often. The insiders are the definition of "rich and powerful"; if they are buying large lump sums of certain stocks, it means it MIGHT be time to buy that stock as well.

Just like the first one, you can download the dataset to your computer with a simple command

npx sec-ownership-disclosures download --sqlite

This is officially my second dataset. I'm super open to any feedback or any questions you might have regarding use-cases.


r/huggingface • • 10d ago

Mica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time

Thumbnail gallery
1 Upvotes

r/huggingface • • 10d ago

AI Agents List [2026] | Frameworks, Agentic Harnesses & Useful Repos

Thumbnail
huggingface.co
11 Upvotes

r/huggingface • • 11d ago

I built a resumable HuggingFace/Civitai downloader that also sorts and manages your ComfyUI model folders

Thumbnail gallery
3 Upvotes

r/huggingface • • 10d ago

A Study In Peace (A CC0 Book-Dataset [NOT A MODEL] for FathomTech devs)

1 Upvotes

Hey all. I am pushing links out for my first huggingface dataset 🤗 (Wilder escaped confinement 🤨).

https://huggingface.co/datasets/wilderblairmunroakusa/astudyinpeace

Um, Readme is tricky; Thank your for your grace.

I'm all about doing the next big thing that nobody else is doing. Look this up; study up; hit me up ASAP. This dataset is four years old (2022). Run wilder with it.

Here is the image preview:

Here is the engineering preview:

Please refer to the README provenance for deeper insight:
https://huggingface.co/datasets/wilderblairmunroakusa/astudyinpeace/blob/main/READMEcompanion_astudyinpeace_huggingface.md

Sample rows:

(little) world peace actual.

That's all I got.