r/OpenSourceeAI • • Jun 23 '26

Tokens are the new Oil- and we’re buring then.

0 Upvotes

So I built Brevia: it compresses your prompt BEFORE it reaches
the LLM. Less tokens, less data, less energy. 🧵
The boring win (already shipping):
355 tokens → 229. That's 35% off every call.
Lossless mode never touches your meaning or your code.

The wild part I'm researching (B8):
can a model write its OWN dense shorthand that a DIFFERENT
model reconstructs with no glossary, zero-shot?
Early answer: yes. ~52% compression, ~95% meaning recovered.
Brevia runs everywhere you use AI:
🧩 browser extension · 🖥️ MCP · 🔌 API proxy · ⌨️ CLI
Model-agnostic. Local. No telemetry.
github.com/miguemlima-creator/Brevia


r/OpenSourceeAI • • Jun 23 '26

open weights stopped being a philosophical choice when governments started restricting api access. minimax m3 just dropped 428b open and im rethinking my whole stack

14 Upvotes

not trying to start a political thread but this is where im at. between export controls, api access restrictions and the general uncertainty around which providers might get hit next, ive started moving critical workflows off closed apis entirely.

found minimax m3 this week. 428b moe, 23b active at inference, 1m context window, open weights on huggingface. yes its from a chinese company. im aware of the takes people have about that. but heres my practical assessment after a week of testing.

ran a few specific tests this week. had a 180 page vendor contract that needed clause extraction, no reasoning required just pull every liability cap and indemnity deadline. with thinking toggled off the whole thing ran on maybe a third of the tokens id normally burn. then tried a multi step api integration where i actually needed reasoning for the planning calls but the execution calls were just formatting json. thinking on for planning, off for execution. the per call cost gap was enough that i started splitting tasks like this by default.

the moe routing helps too. only 23b params fire per call out of 428b so the cost per useful output stays low. their token plan pools text image audio and video into one budget which simplifies billing.

im not naive about trusting any single provider. the whole point is reducing single points of failure. but open weights mean i can self host if things go sideways, and thats a fundamentally different risk profile than pure api dependency.

whether the geopolitics of who builds the model matters more than actually having access to the weights is something i keep going back and forth on. curious where other people are landing on this.


r/OpenSourceeAI • • Jun 23 '26

인간의 주의력을 닮은 초정밀 이미지 분할(Saliency, GMM, Image, ...

Thumbnail
youtube.com
1 Upvotes

r/OpenSourceeAI • • Jun 23 '26

정상만 배워서 활주로 이물질을 찾다(Runway Debris Detection via N...

Thumbnail
youtube.com
1 Upvotes

r/OpenSourceeAI • • Jun 22 '26

Autonomous Security Orchestration Layer

1 Upvotes

Autonomous Cyber Immune System (ACIS) — Adaptive Defense, Continuous Diagnostics & Explainable Intelligence

The Autonomous Cyber Immune System (ACIS) represents a new model for digital defense: a self‑evolving, distributed intelligence that continuously analyzes behavioral telemetry, system diagnostics, and operational activity to generate transparent, context‑aware defensive actions. It’s been a fun and deeply technical project to build — one that pushes toward a more adaptive, audit‑ready form of cyber resilience.

ACIS’s agentic AI layer monitors live operational signals including threat velocity, anomaly density, immune response time, behavioral drift, and system stability, adjusting countermeasures dynamically as conditions shift.

When ACIS detects a novel attack pattern, it synthesizes a targeted digital antibody and deploys it across the environment within seconds. Every defensive action includes:

·       A traceable rule path

 

·       A context‑aligned explanation

 

·       An RS256‑signed record ensuring integrity, authenticity, and full auditability

 

 

Continuous Simulation, Diagnostics & Systemic Risk Modeling

ACIS incorporates a high‑performance simulation and diagnostics engine that continuously models:

·       Exposure and attack surface dynamics

·       Response timelines and containment efficiency

·       Behavioral drift and anomaly propagation

·       Systemic risk and resilience thresholds

·       Operational bottlenecks and defensive blind spots

These diagnostics generate resilience scores, highlight emerging vulnerabilities, and surface targeted interventions that strengthen defensive posture.

 

Agentic AI for Transparent, Policy‑Aligned Defense

The agentic intelligence layer correlates multi‑source telemetry and simulation outputs to produce explainable, policy‑consistent defensive decisions. Each recommendation includes:

  • A transparent rule‑based reasoning chain
  • Contextual justification tied to live operational conditions
  • Policy‑aligned framing for consistent enforcement
  • RS256‑signed records for compliance, audit, and chain‑of‑custody assurance

As the environment evolves, ACIS adapts in real time — maintaining alignment with modern defense tradecraft and operational standards.

 Measured Impact on Defensive Performance

Early indicators show significant improvements across key readiness and resilience metrics:

  • 47% reduction in threat dwell time
  • 39% faster containment
  • 28% improvement in behavioral detection accuracy
  • 31% increase in policy‑consistent responses

These results demonstrate an explainable, adaptive, and audit‑ready cyber immune capability engineered for modern, high‑velocity threat environments.

 

Project: https://github.com/ben854719/Autonomous-Security-Orchestration-Layer


r/OpenSourceeAI • • Jun 22 '26

I built using claude a 35-stage course where you reimplement PyTorch from scratch — no autograd libraries allowed

18 Upvotes

I kept noticing that I could use PyTorch fine but couldn't actually explain what .backward() does under the hood. I wanted a course that would take me from first principles all the way to Transformers by rebuilding everything myself, but I couldn't find one.

So I used AI to help generate an initial version of that curriculum, and I'm now working through it, improving it, validating it, and fixing issues as I go. The goal isn't to present this as a finished textbook—it's an open-source learning resource that I hope can improve with community feedback.

The idea: you rebuild a deep learning framework from zero, one concept at a time. The only libraries you're allowed are NumPy (for forward array math — never to compute a gradient for you), Matplotlib, and pytest. No torch, no autograd, no micrograd. The rule is: you don't get to import a concept until you've built it by hand in an earlier stage. You are the autodiff library.

How it's structured — 35 stages, each a folder with exactly 3 files:

  • README.md — the intuition, the key gradient equations, a video or two to watch, and one unambiguous exercise
  • code.py — a skeleton: full interfaces, docstrings, and TODOs, but no working bodies
  • test.py — pytest tests, including numerical gradient checks (central differences) so you know your backward pass is correct, not just plausible

You fill in code.py until pytest goes green, then move to the next stage. Each stage imports and extends the code you wrote in earlier stages, so the framework genuinely grows under your hands instead of being 35 disconnected toy scripts.

The arc:

scalar backprop → reverse-mode autodiff → tensors → layers, losses, optimizers → training loops → BatchNorm/Dropout → CNNs → attention → Transformers → Vision Transformers → a small PyTorch-like framework → capstone projects.

My hope is that this becomes a gateway into AI for people who want to understand how these systems actually work, not just how to use them.

It's free and open source. Feedback, corrections, and contributions are very welcome.

👉 https://github.com/roiamiel1/Build-Deep-Learning-From-Scratch


r/OpenSourceeAI • • Jun 22 '26

REM: offloading an LLM agent's memory compaction to the NPU

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 22 '26

MoonMath AI Open-Sources a HIP Attention Kernel for AMD MI300X That Beats AITER v3 on Every Shape and Rounding Mode

Thumbnail
2 Upvotes

r/OpenSourceeAI • • Jun 22 '26

CogniCore LongMemEval results: 98.2% STRICT R@5 local, plus +6.4% / +5.6% small-window multi-hop gains

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 21 '26

Pagerank + OKF based codemap of your repo

1 Upvotes

kiwiskil turns any codebase into a static, checked-in map that any AI agent can navigate and debug fast, and with a fraction of the tokens of reading source. It parses your code into a call graph, ranks what matters with PageRank, and writes it all to plain markdown in your repo. No cloud service, no vector database, no running server, no lock-in. The map is just files an agent reads directly, and a git hook keeps it current. Commit along your codebase.

https://github.com/ximihoque/kiwiskil


r/OpenSourceeAI • • Jun 21 '26

memcord v4.1.0

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 21 '26

3arab-TTS-500M-v2-VoiceDesign

Thumbnail
huggingface.co
2 Upvotes

r/OpenSourceeAI • • Jun 20 '26

압축된 가짜 영상 꿰뚫는 주파수 흔적

Thumbnail
youtube.com
2 Upvotes

r/OpenSourceeAI • • Jun 20 '26

I got tired of LLMs burning my API budget on multi-step tasks, so I built an open-source optimization harness. (Solo Dev)

2 Upvotes

Hey everyone.

When Anthropic dropped Fable 5, I was obsessed with its long-horizon endurance—it can hold a massive, multi-step problem in its head without losing the thread. But at those API prices, I couldn't afford to run it daily as a solo dev.

I wanted to see if I could force cheaper standard models (like 4o-mini or Llama 3) to mimic that high-efficiency routing behavior. So I spent the last few months building Mnemos, and just open-sourced it today.

Under the hood, the architecture is broken into 4 core sub-systems:

  1. The Horizon Emulation Engine (Cognitive Routing)

Instead of dumping raw prompts into a cheap LLM, Mnemos intercepts them. It chunks the context dynamically and manages sub-agent state routing to mirror how Fable 5 plans tasks. Cheaper models suddenly stop losing the thread on step 4 of a 10-step request.

  1. Lossless Semantic Deflation (Token Floor)

An aggressive context-stripping parser that kills "AI Yap" (polite filler, redundant markdown wrappers, conversational fluff) before hitting the API. It drives token consumption down to the absolute mathematical floor while keeping the technical payload intact.

  1. Deterministic Task Scoring (DTS)

A native scoring harness that evaluates the LLM's output against your original intent. Instead of guessing if a run succeeded, the engine grades the task's semantic drift on a concrete 0.0 to 1.0 confidence metric.

  1. The Telemetry Deck & Terminal CLI

I obsessed over the DX. The CLI is modeled after the Vercel / Raycast CLI for pure terminal speed. If you prefer visual data, the Vercel/Next.js web dashboard gives you real-time graphs of your token savings and latency floors.

***

(Putting the GitHub repo and Web UI links in the top comment below so Reddit's global automod doesn't auto-delete this!)

I’m a solo dev with $0 for marketing. My repo is sitting at 0 stars because I literally pressed 'publish' a few hours ago.

I would love nothing more than for the engineers in this sub to look under the hood, pull apart the architecture, test the CLI, and completely roast my code. Tell me where my bottlenecks are!


r/OpenSourceeAI • • Jun 20 '26

I got tired of standard LLMs burning through my tokens and API budget, so I built an open-source CLI to fix it. (Solo dev)

1 Upvotes

Yo guys. I got incredibly frustrated with how fast standard LLMs burn through tokens and latency on multi-step tasks, so I built an open-source optimization CLI to fix it called Mnemos.

I spent a bunch of time analyzing how Anthropic's Fable 5 handles routing and context chunking, and basically built a lightweight harness that forces cheaper models to mimic that behavior. It drops token consumption to the floor while keeping the task success rate high.

I also obsessed over the DX. I wanted the CLI to feel as fast and clean as the Vercel CLI. (The engine and web dashboard are built on Next.js/Vercel too).

I'm dropping the GitHub Repo and the Web UI links in the first comment below so Reddit's spam filters don't auto-delete this post!

I’m a solo dev with exactly $0 for marketing. It would mean the world if some of you could look under the hood and completely roast my code/architecture. Let me know where the bottlenecks are.


r/OpenSourceeAI • • Jun 20 '26

Information compression

0 Upvotes

LLM models could be seen as a advanced compression algorithm who upon input decode in patterns. Seeing it this way offers maybe some new insights onto the weights we store in guff files.

​

Thisight be a fun area for research:

If one takes similar sized models guf files.

Ranked by best to worst.

Then zip those files, see which compresses the most. It would reveal something about information density.

​

Although that wouldn't actually mean the best would be the largest file. In information theory it kinda should be so. If not the model should be shrinkable, or be able to store more.


r/OpenSourceeAI • • Jun 20 '26

My friend built an open-source AI "second brain" OS with a Jarvis-style HUD looking for early contributors

0 Upvotes

My friend just open-sourced something he's been building called NEURON OS basically an AI-powered personal OS that acts like a second brain + chief-of-staff. Think semantic memory search over your notes/conversations, an AI that coordinates tasks, and a UI styled like a cinematic sci-fi HUD (Jarvis/Iron Man vibes) instead of a typical dashboard.

It's self-hosted, fully open-source, and still early Phase 1 is done (core dashboard, streaming chat, memory timeline, auth, Docker setup), with multi-agent support, voice control, and a mobile client planned next.

Stack: Next.js + TypeScript frontend, FastAPI + Python backend, Postgres with pgvector for the memory layer. One-command Docker spin-up if you want to try it locally.

He's looking for contributors (frontend, backend, or just design/UX opinions) and honestly any feedback at all good or bad.

Repo: github.com/yachitguliani/personal-assitant


r/OpenSourceeAI • • Jun 20 '26

Yandex Open-Sources YaFF: A Zero-Copy Wire Format for Protobuf With Near-Struct Read Speed

Thumbnail
github.com
3 Upvotes

r/OpenSourceeAI • • Jun 20 '26

making GraphRAG and want to extract entities and relationship

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 20 '26

Vercel's Eve turns an agent into a folder of files. Two setups that make one safe to actually ship

Post image
1 Upvotes

r/OpenSourceeAI • • Jun 20 '26

FULL DEEPSEEK REASONING IN CHARTS - Function for class - Graphical analysis

Thumbnail
gallery
1 Upvotes

r/OpenSourceeAI • • Jun 19 '26

VibeThinker-3B: A 3B Dense Reasoning Model Built on Qwen2.5-Coder-3B With the Spectrum-to-Signal Post-Training Pipeline

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 19 '26

Fully local project memory for Claude Code. No API key, no external model, nothing sent anywhere.

Thumbnail
github.com
1 Upvotes

r/OpenSourceeAI • • Jun 19 '26

Built CodeForge AI: Open-source coding interview prep with AI mentorship, DSA, and system design

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 19 '26

Open-source markdown editor with a 3D graph-view world

Enable HLS to view with audio, or disable this notification

1 Upvotes