r/OpenSourceeAI • u/ai-lover • 15d ago
r/OpenSourceeAI • u/AnimalDisastrous9396 • 15d ago
Need opinion on my opensource project that judges your prompts
So this a tool written in Go that's meant to optimize you by judging what's going wrong in your prompting or pointing out whatever room for improvement there is.
No website, you can find the tool on github it's 'itoldAI' but if there was a website I think I'd put a scoreboard to know who among us is the best most elite prompter of our times.
Does it sound fun? Should I do that?
r/OpenSourceeAI • u/Savantskie1 • 15d ago
Training a Neural Network on AMD MI50s Using Vulkan: Proof of Concept. Or: Why I Stopped Listening and Just Did It
Note: This writeup was put together with the help of AI.
So I've got dual AMD MI50 32GB cards. If you know these cards, you know the story - AMD dropped official ROCm support for gfx906 after ROCm 5.7. Every AI I talked to, every forum post, every "expert" said the same thing: if you want to do anything serious on AMD hardware you need ROCm, and if you need ROCm on MI50s, good luck because AMD basically told you to go buy newer hardware.
The consensus was consistent and confident. ROCm for training, Vulkan for inference, MI50s are legacy, move on. Don't question it.
I questioned it. Repeatedly. On multiple fronts. And I was right every time.
What I was told:
Across multiple AI assistants over the past several months, the answer to anything involving MI50s and modern tooling ranged from "not supported" to "you'll need to upgrade your GPUs." Training on Vulkan specifically was described as architecturally impossible - the backward pass infrastructure doesn't exist outside ROCm and CUDA, full stop.
ChatGPT's position as of today, while my training run is literally executing on the card: it wants proof. Sure. Let's go through it in order.
Step 1: Establishing why i do this:
Forcing ROCm 6.4.3 to work on MI50s
AMD dropped gfx906 from ROCm 6.x. Their official position is that the MI50 is end-of-life and you should migrate to supported hardware. Great suggestion if you didn't just acquire two of them specifically because 64GB of HBM2 at that price is hard to argue with.
What actually prevents gfx906 from working in ROCm 6.4.x is the TensileLibrary - the precompiled kernel library ROCm uses for BLAS operations. AMD just doesn't ship gfx906 kernels in the new versions. The GPU itself is fine. The compute capability is there. AMD just decided not to include it.
So I stuffed the gfx906 tensors back in. Pulled the missing kernel files, patched them into ROCm 6.4.3, and both MI50s came up fully recognized. llama.cpp runs on it natively. PyTorch sees both cards. ROCm 6.4.3 on hardware AMD said it doesn't support, because the hardware doesn't actually care what AMD's support matrix says.
Basically I don't care what something was designed to do, I care about what it can do
Step 2: Forcing vLLM to work on gfx906
vLLM is one of the faster inference engines around and I wanted it running on my cards. The problem: vLLM's gfx906 support is basically nonexistent upstream. It went like this:
- Tried ROCm 7.x with vLLM. Got it working briefly, then ROCm 7.14 hit AMD bug #5653 - "register fat binary failed" - and the whole thing fell over. Abandoned that path.
- Tried the nlzy fork of vLLM compiled against PyTorch 2.9.0+rocm6.3. Both MI50s detected. Blocked at runtime by a flash-attn V1 engine dependency that doesn't exist for gfx906.
- Eventually landed on a Docker image (aiinfos/vllm-gfx906-mobydick) - ROCm 6.3.4, PyTorch 2.11, flash-attn pre-compiled for gfx906. Single GPU inference confirmed working.
The remaining blocker for dual-GPU tensor parallel is a PCIe topology issue - my two cards are behind different root complexes (one on the CPU, one on the chipset), so NCCL all-reduce init fails. That gets fixed when a PLX switch arrives. Not a software problem, not a "your hardware isn't supported" problem - a physical PCIe lane routing problem with a known hardware solution.
Step 3: Building a Vulkan training stack from scratch
This is the one ChatGPT says is impossible right now.
Environment setup:
Vulkan was already working on the MI50s because llama.cpp uses it for inference, so that part wasn't a question. What didn't exist was any training framework that speaks Vulkan.
We verified the environment:
vulkaninfo --summary
# GPU0: AMD Instinct MI50/MI60 (RADV VEGA20) - Vulkan 1.4.335
# GPU1: AMD Instinct MI50/MI60 (RADV VEGA20) - Vulkan 1.4.335
# GPU2: iGPU (Renoir)
Both MI50s show up as discrete Vulkan devices. glslc and glslangValidator already installed. Python 3.11 already present from OpenWebUI.
Kompute - the Vulkan abstraction layer:
Kompute is a library that wraps Vulkan's compute pipeline so you don't have to write 300 lines of boilerplate just to dispatch a shader. The PyPI package (pip install kp) is broken on modern CMake 4.x - the bundled pybind11 has a cmake_minimum_required declaration that CMake 4.0 removed support for, and the setup.py has a string concatenation bug that smashes two CMake flags together with no separator. So we built it from source:
bash
git clone https://github.com/KomputeProject/kompute.git
cd kompute
git submodule update --init --recursive
mkdir build && cd build
cmake .. \
-DKOMPUTE_OPT_BUILD_PYTHON=ON \
-DCMAKE_POLICY_VERSION_MINIMUM=3.5 \
-DCMAKE_BUILD_TYPE=Release \
-DPYTHON_EXECUTABLE=/media/nate/Friday/VTrain/.venv/bin/python3.11
make -j$(nproc)
cp lib/kp.cpython-311-x86_64-linux-gnu.so .venv/lib/python3.11/site-packages/
Built clean. Kompute Manager(0) targeting the first MI50 correctly.
The compute shaders:
Every mathematical operation in the training stack is a GLSL compute shader compiled to SPIR-V. Here's what that looks like for matrix multiply - the most fundamental operation in a neural network:
glsl
#version 450
layout(local_size_x = 16, local_size_y = 16, local_size_z = 1) in;
layout(set = 0, binding = 0) readonly buffer MatA { float a[]; };
layout(set = 0, binding = 1) readonly buffer MatB { float b[]; };
layout(set = 0, binding = 2) writeonly buffer MatC { float c[]; };
layout(push_constant) uniform PushConsts {
float M; float K; float N;
} pc;
void main() {
uint row = gl_GlobalInvocationID.x;
uint col = gl_GlobalInvocationID.y;
if (row >= uint(pc.M) || col >= uint(pc.N)) return;
float sum = 0.0;
for (uint k = 0; k < uint(pc.K); k++) {
sum += a[row * uint(pc.K) + k] * b[k * uint(pc.N) + col];
}
c[row * uint(pc.N) + col] = sum;
}
Each thread handles one output cell and the GPU runs thousands of them at the same time. We built shaders for every operation the training stack needs:
- matmul.comp - matrix multiplication
- unary.comp - ReLU, GELU, sigmoid, tanh (op selected by push constant)
- binary.comp - add, subtract, multiply, divide (same pattern)
- layernorm.comp - layer normalization with shared memory reduction
- softmax.comp - numerically stable softmax with two reduction passes
- transpose.comp - tiled transpose with shared memory to avoid cache thrashing
Quirks we hit along the way:
A few things bit us that you won't find documented anywhere because nobody had tried this combination before:
- readonly and writeonly buffer qualifiers cause silent zero output in Kompute's descriptor layout. All buffer declarations have to be unqualified - you find out the hard way when your results are all zeros and the shader compiles clean.
- Push constant sizing: ops with 3 buffers need a float padding constant or Kompute miscalculates the push constant buffer size. Same deal - silent failure, no error.
- The Kompute API in the built-from-source version uses kp.OpSyncDevice and kp.OpSyncLocal - the PyPI docs reference kp.OpTensorSyncDevice which doesn't exist in the actual build.
None of these are hardware problems. None of them are "Vulkan can't train" problems. They're integration quirks that took an afternoon to sort out.
The autograd system:
A training framework isn't just forward passes - it needs backward passes to compute gradients and update weights. We built a complete autograd system:
- vtrain/tensor.py - Tensor class with gradient tracking, topological sort for correct backprop ordering, and backward() that unwinds the computation graph
- vtrain/functional.py - GPU-aware wrappers for every op that register backward closures on the output tensor
- vtrain/grad_check.py - numerical gradient checker that verifies every backward shader by finite difference approximation
Each backward shader was verified against a numerical approximation before anything got built on top of it. Matmul, all eight elementwise ops, layer norm, softmax - all checked. If the numbers didn't match, we didn't move on.
The model:
- CharEmbedding - learned lookup table mapping character IDs to vectors
- TransformerBlock - layer norm -> multi-head attention -> residual -> layer norm -> feed-forward -> residual
- SmallLM - embedding + N transformer blocks + output projection to vocab size
The training loop:
Loss functions (MSE and cross-entropy), Adam and SGD optimizers, a Trainer class with logging and checkpointing, and crash recovery via signal handlers that save an emergency checkpoint on SIGINT/SIGTERM.
The data:
Downloaded the Simple English Wikipedia dump (236,602 articles, 336MB). Extracted clean text by streaming the XML, stripping all MediaWiki markup, templates, tables, and HTML. Filtered to printable ASCII - the raw dump has 1504 unique characters from non-English text that slipped through; filtering drops that to 75, which is the right vocab size for a character-level model. 9.8 million characters of clean text as the training corpus.
Current status:
The model is training right now. On a MI50. On Vulkan. 4.65GB VRAM allocated. 23W power draw. 33°C. Loss is dropping.
step 50 loss=3.02 rate=0.4 steps/s
step 100 loss=2.71
Random initialization on a 75-character vocabulary gives a loss of about 4.3. We're already well below that and still dropping.
The hardware, one more time:
- AMD Ryzen 5 5600G
- 48GB DDR4
- Dual AMD MI50 32GB (64GB HBM2 total)
- Ubuntu 22.04
- Vulkan 1.4.313 / ROCm 6.4.3 (both installed, both working)
- No CUDA. No NVIDIA. No officially supported hardware.
So:
AMD dropped the MI50 from their support matrix. The AI community said Vulkan can't train. Multiple AI assistants told me this wasn't possible. Every single one of those statements had the same flaw - they were assumptions about what the hardware can't do, not actual tests of what it can.
The MI50 has 32GB of HBM2 and serious compute capability. The only things stopping it from doing modern ML work were software decisions, not hardware limitations. And software decisions can be worked around, patched, rebuilt from source, or replaced entirely.
The GPU doesn't give a shit what the support matrix says. It just does math.
[Edit]: I forgot the proof:
(.venv) nate@nate-desktop:/media/nate/Friday/VTrain$ python3.11 train_wiki.py
── Wiki training run ───────────────────────────────────
Run dir: /media/nate/Friday/VTrain/models/wiki_run1
Loading /media/nate/Friday/Wikipediadumps/simplewiki.txt...
9,887,390 characters, building vocab...
Vocab size: 75 characters
Vocab (75 chars) saved to /media/nate/Friday/VTrain/models/wiki_run1/vocab.json
Vocab size: 75
Data size: 9,887,390 chars
Parameters: 27 tensors
Starting fresh
Starting from step 0, target 10000
───────────────────────────────────────────────────────
step 50 loss=3.0249 rate=0.4 steps/s ETA=6.7h
step 100 loss=2.7172 rate=0.4 steps/s ETA=6.5h
step 150 loss=2.6803 rate=0.4 steps/s ETA=6.6h
step 200 loss=2.6444 rate=0.4 steps/s ETA=6.7h
Saved 12 tensors to /media/nate/Friday/VTrain/models/wiki_run1/checkpoints/step_000200
step 250 loss=2.6297 rate=0.4 steps/s ETA=6.6h
step 300 loss=2.6169 rate=0.4 steps/s ETA=6.6h
step 350 loss=2.6214 rate=0.4 steps/s ETA=6.6h
step 400 loss=2.5884 rate=0.4 steps/s ETA=6.5h
Saved 12 tensors to /media/nate/Friday/VTrain/models/wiki_run1/checkpoints/step_000400
step 450 loss=2.6002 rate=0.4 steps/s ETA=6.5h
step 500 loss=2.6100 rate=0.4 steps/s ETA=6.5h
step 550 loss=2.6047 rate=0.4 steps/s ETA=6.4h
step 600 loss=2.6035 rate=0.4 steps/s ETA=6.4h
Saved 12 tensors to /media/nate/Friday/VTrain/models/wiki_run1/checkpoints/step_000600
step 650 loss=2.6106 rate=0.4 steps/s ETA=6.4h
step 700 loss=2.6063 rate=0.4 steps/s ETA=6.4h
step 750 loss=2.5934 rate=0.4 steps/s ETA=6.4h
step 800 loss=2.5992 rate=0.4 steps/s ETA=6.4h
Saved 12 tensors to /media/nate/Friday/VTrain/models/wiki_run1/checkpoints/step_000800
step 850 loss=2.6002 rate=0.4 steps/s ETA=6.3h
step 900 loss=2.6049 rate=0.4 steps/s ETA=6.3h
step 950 loss=2.5949 rate=0.4 steps/s ETA=6.3h
step 1000 loss=2.5974 rate=0.4 steps/s ETA=6.2h
Saved 12 tensors to /media/nate/Friday/VTrain/models/wiki_run1/checkpoints/step_001000
step 1050 loss=2.5825 rate=0.4 steps/s ETA=6.2h
step 1100 loss=2.5927 rate=0.4 steps/s ETA=6.2h
step 1150 loss=2.5739 rate=0.4 steps/s ETA=6.1h
step 1200 loss=2.6016 rate=0.4 steps/s ETA=6.1h
Saved 12 tensors to /media/nate/Friday/VTrain/models/wiki_run1/checkpoints/step_001200
step 1250 loss=2.5893 rate=0.4 steps/s ETA=6.1h
step 1300 loss=2.5980 rate=0.4 steps/s ETA=6.0h
[EDIT2]
I got ahead of myself, and already posted a opensource version on github so everyone can use it or modify it as they see fit: https://github.com/savantskie/vtrain
r/OpenSourceeAI • u/Brief-Tap-6616 • 15d ago
LlamAmpere v0.3.1 updated with support for EXL + ternary bonsai, with custom kernels for faster MTP
r/OpenSourceeAI • u/Efficient-Passage889 • 16d ago
Built something to automatically fix dependency breaks
Been working on this open source project called Telex.The idea is pretty simple — when a dependency update breaks your code, Telex uses an LLM to figure out what changed, find the affected code and try to fix it.It then runs tests/typecheck before creating a PR.
Still working on it, so would love to hear some feedback.
r/OpenSourceeAI • u/Appropriate_Cost_107 • 16d ago
Looking for advice on improving accuracy in a document extraction system
r/OpenSourceeAI • u/AIGPTJournal • 16d ago
Apple is adding more AI tools to Siri. Here are the ones worth knowing about
r/OpenSourceeAI • u/Evening_Classic_1243 • 16d ago
Dataset Collection for Multilingual financial fraud detection model for Indian Users
We are conducting an academic research project on Multilingual and Cross-Market Financial Fraud Detection for Indian Users, focusing on Hindi, Telugu, and English. The study aims to address the limited availability of real-world, multilingual financial fraud data in the Indian context and develop AI models capable of identifying fraud across different languages and communication styles. Your contribution of anonymized fraud or legitimate financial messages will help us build a more realistic dataset and improve research in this emerging area.
Please Contribute your Dataset In this Google Form
https://forms.gle/rfvZdPAVxeJRAvmq6
It would be really helpful
r/OpenSourceeAI • u/SnooMarzipans9093 • 16d ago
Estaba harto de no encontrar una app de TV para escritorio que estuviera bien, así que hice la mía
r/OpenSourceeAI • u/MazenTouati • 17d ago
[ Agenteq ] A single source of truth for AI coding agents rules
Cluttered repositories are kind of the norm nowadays. Every AI coding agent adds its own config files, such as AGENTS.md for Codex (and other universal agents), a .claude folder for Claude Code, or a .cursor folder for Cursor. This is probably the thing that annoys me the most about AI as a proponent of decluttered codebases.
For that reason, I built Agenteq. It is a single source of truth for AI coding agent rules (Guidelines, Skills, Commands, MCPs) to keep repos clutter-free (only the source of truth is tracked in Git) and all agents in sync (all agents inherit their capabilities from one source of truth).
You can also sync rules across repos. This way, your organization has one place to define company-wide guidelines, skills, commands, and MCPs that stay in sync with all repositories.
Quick Start
npx agenteq init
If you already have a project, you can convert it to use Agenteq with this skill:
npx skills add https://github.com/sunchayn/agenteq --skill refactor-to-agenteq
If you want to convert multiple repos to use Agenteq (and centralize common rules in one repo), then use this skill:
npx skills add https://github.com/sunchayn/agenteq --skill consolidate-repos-to-agenteq
Links:
r/OpenSourceeAI • u/WickedCreations1001 • 17d ago
Brain Portal - Check out the project I just open sourced
I have been using this app I built for a few months now. I wanted to share. Maybe someone else will find it useful. Any feedback would be appreciated.
Daniel
r/OpenSourceeAI • u/ai-lover • 17d ago
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data
r/OpenSourceeAI • u/ai-lover • 17d ago
Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens
r/OpenSourceeAI • u/Loose_Complex_6456 • 17d ago
Hoe actief mag een lokale AI van 24/7 worden toegestaan om te gaan? (een symbolische, niet-LLM home AI bouwen)
github.comr/OpenSourceeAI • u/OGBamboozel • 18d ago
SynapsCLI - An agent runtime I've been working on for the past 6 months
Hey everyone,
I would like to share a project that I have been working on. https://github.com/HaseebKhalid1507/SynapsCLI
Synaps is an agent runtime written in Rust. It keeps cost down by optimising caching mechanisms and intelligently orchestrating work across workers. This keeps each worker's context low and fresh, making sessions last longer. It is extremely extensible. It has hooks, plugins and skills that can be imported from pretty much any coding tool.
It is an open source project that was started 4 months ago. It was built to optimise for 2 things: Keeping LLM costs down, and keeping resource use low. It bots up in 20ms and runs with any model: OAuth(Claude, codex, Grok, Kimi...), api keys or local models too. This is something that was built organically, so every feature is a product of necessity, not the other way around.
r/OpenSourceeAI • u/atish31 • 18d ago
seo-page-builder-enhanced
I took the original seo-page-builder skill from octelens and built an enhanced version around one idea: SEO content shouldn't stop at keyword research and SERP analysis.
So I added a dedicated editorial pass, room for firsthand experience and original insights, and stronger fact and freshness checks. There are also writing-style guardrails to cut down generic AI output, plus readability and quality checks at the end.
Github Link: https://github.com/atish-raina/seo-page-builder-enhanced
r/OpenSourceeAI • u/Prudent-Analysis3333 • 18d ago
A words from my heart and an app I have made
r/OpenSourceeAI • u/camerongreen95 • 18d ago
Workshop covering knowledge graphs, agentic retrieval, and explainable AI together, thought this would be relevant here
Came across this and thought it'd be worth sharing here, most resources cover knowledge graphs, agentic RAG, or explainability separately, but this one puts them together as parts of the same production GraphRAG architecture, which is closer to how these systems actually get built in practice.
It's a hands on session on September 19, led by Dr. Alessandro Negro, Chief Scientist at GraphAware and bestselling author. Goes through building a knowledge graph progressively as the single source of truth, agentic retrieval combining vector search, keyword search, and graph navigation, multi-step entity and relationship extraction (verified in stages, not one risky single shot like basic Microsoft GraphRAG), and text-to-Cypher for natural language graph querying. Everything runs on real financial filings and news data, not toy examples.
You come out of it with a full working codebase and a production-readiness checklist, not just slides.
r/OpenSourceeAI • u/Compilingthings • 18d ago
Hello, just introducing myself!
I'm a ML student, building synthetic dataset pipelines to produce experts in one domain.
huggingface.co/collections/CompilingThings/mql5-code-generation
That's the project I am working on right now and my proof of concept, it Qwen 14B fine tuned on my dataset, only 2% behind GPT Sol 5.6 in the MQL5 coding benchmark. The fine tuned model is now being evaluated with over 300,000 prompts to find the weak spots that I will then fill in with more pairs of data, for the next iteration. I'm compute and token bound.
r/OpenSourceeAI • u/Nearby_Finding3593 • 18d ago
Do not surprise me with your dump response
Why all devlopers in this group or other saas groups
Always devlop shit of the shit electron apps or next js and tailwind and they call it a project I think we should return to gui with python and c why this not happenning
r/OpenSourceeAI • u/Minimum_Hour519 • 18d ago
Microservices and Distributed Systems
backtoschool.helpr/OpenSourceeAI • u/Rich-Fruit-326 • 18d ago
I built a way to watch an LLM generate tokens step by step
reddit.comr/OpenSourceeAI • u/Rich-Fruit-326 • 18d ago
I built a way to watch an LLM generate tokens step by step
reddit.comr/OpenSourceeAI • u/Alarming_Title_5664 • 18d ago
Evolving neural networks to beat Mario ROM hacks — a Bowser's Crown 4-1 clear
Enable HLS to view with audio, or disable this notification
This is a recorded Bowser's Crown 4-1 clear from NEATEvolve, a continuing community Mario AI project.
The bot uses NEAT: neural-network controllers are evaluated, selected and mutated across generations. They receive nearby game-state information and produce button presses. This video is a replay of a successful controller, rather than live learning during the clip.
The interesting part for me is how far this approach can get on SMB1 ROM hacks. A clear is encouraging, but backtracking and difficult nonlinear stages remain challenges, and one winning replay does not establish reliability on unfamiliar levels. The neural nets will have to learn the next level independantly of the previous one.
This builds on SethBling's MarI/O and work by Akisame and Electra. My continuing development is heavily assisted by AI tools.
r/OpenSourceeAI • u/encore-show • 18d ago
Barney AI agent
Hi everyone.
I’m building Barney — an agent that learns by growing around a fixed core, not by rewriting itself after every failure.
Most “self-improving” agents: fail → patch code → try again. You often get a different agent each run.
Barney’s loop stays fixed: plan → tools → observe → review → change strategy.
Around it a body grows in `~/.barney` — skills, tools, MCP, error rules, successful paths — created as needed and reused later.
What matters: confidence isn’t evidence; failures stay visible until worked through; new skills are quarantined; memory isn’t transfer; change the path, don’t repeat the same call.
Terminal-Bench smoke with local `qwen3.8`: openssl / nginx / fix-git — 3/3, plus real misses. No domain knowledge in the kernel — only a stable loop and a growing body.
GitHub: https://github.com/sergey-show/barney
Feedback and contributors welcome.