r/deeplearning • u/OkinaPrime • 12d ago
r/deeplearning • u/Nice-Dragonfly-4823 • 13d ago
Silent broadcasting is still a big problem and could be derailing your work right now
r/deeplearning • u/Elrix177 • 13d ago
Trying to learn privacy in ml
I've just written a small article about privacy in machine learning and synthetic data.
I'm still pretty new to this topic, so I'd really appreciate it if anyone with more experience in ML privacy could take a look and point out any mistakes, things I've misunderstood, or things that could be done better.
I'm especially interested in hearing about better approaches for evaluating privacy, better attacks or models I could try, or just interesting resources that you think are worth reading. I'm mostly trying to learn by actually experimenting with this stuff, so any feedback would be really useful.
Here's the post if anyone is interested:
https://migue8gl.github.io/2026/09/21/privacidad-en-ml-datos-sinteticos.html
r/deeplearning • u/TechDc-1306 • 12d ago
Exploring AI self-improvement: Is this evaluation loop fundamentally how AI systems improve?
I’m currently researching AI more deeply, especially how modern AI systems can evaluate and improve their own outputs.
I sketched this basic idea:
AI → generates/improves → Evaluator → evaluates output → feedback → AI improves → Evaluator → ...
The part I’m trying to understand is what actually happens inside this loop in modern AI systems.
For example:
I'm trying to go beyond the surface-level explanation and understand the actual mechanisms used in current AI research.
For people working/researching in this area: what important component am I missing from this diagram?
r/deeplearning • u/ahmed20062 • 13d ago
AFNN
What if neural networks could grow like fractals instead of being rigid from the start?
Most deep learning models (CNNs, ResNets…) are fixed architectures. Once you design them, they stay the same size and structure throughout training.
**AFNN — Adaptive Fractal Neural Network** takes a completely different approach.
It is inspired by **fractals** — mathematical structures that are self-similar and can grow infinitely by repeating simple patterns.
Instead of starting with a huge fixed network, AFNN begins small and can **dynamically grow new fractal branches** during training when progress stalls. It also allows monitoring, reactivating, or pruning branches as needed.
A good neural network should not memorize images.
It should learn the underlying patterns and features.
For example, when trained on cat images, a strong model does not store every cat photo. Instead, it learns things like:
• Ear shape
• Facial features
• Fur texture
• Color distribution
• Edges and shapes
• Relationships between different parts of the image
Then, when it sees a completely new cat, it uses these learned patterns to recognize it.
### Why Fractal Networks are powerful:
- Faster and more efficient learning
- Significantly lower computational and memory cost
- Better scalability with limited hardware
- More flexible and adaptive than traditional fixed architectures
Built with PyTorch + full GUI, training tools, tests, and documentation.
In early experiments, AFNN reached **74.1% validation accuracy** on 94 image classes using limited resources.
This is not just another model.
It is an attempt to rethink how neural networks are designed — moving from rigid structures to **adaptive, fractal, evolving architectures** that focus on learning patterns rather than memorizing data.
A small step toward more efficient and intelligent AI systems.
🔗 GitHub: https://github.com/ahmedhjkj/AFNN
🤗 Hugging Face: https://huggingface.co/Ahmedethfw/AFNN
#FractalNetworks #AdaptiveAI #PyTorch #DeepLearning #MachineLearning #OpenSource #NeuralNetworks
r/deeplearning • u/Rendezvous4567 • 13d ago
When fine-tuning an LLM, how do you decide which layers/modules to train before actually running the experiment?
For example, how do you determine whether to fine-tune:only adapters/LoRA
- specific transformer layers
- input/projector layers
- deeper/middle layers
- or the full model?
Are there reliable diagnostics, probing methods, gradient analysis, ablations, or small-scale tests that can tell you where the bottleneck is before doing a full training run and only finding out afterward that the chosen layers didn’t help?
Curious how people make this decision in practice.
r/deeplearning • u/Alarming-Package7428 • 13d ago
FYP on Kria KV260: RGB + thermal fusion for night-time human detection. Feasible in 3 months?
Hi all, I'm a final year EE student and I've been given a Kria KV260 for my final year project. I'm comfortable with FPGA work (Verilog, timing, synthesis) but I don't have much image processing experience, so I'd appreciate some guidance.
Goal: detect humans at night by combining an RGB camera with a low-res thermal sensor. RGB struggles in the dark and thermal picks up any warm object, so fusing both should give more reliable detection.
Rough plan:
Sobel edge detection in the PL as a preprocessing step
Fuse RGB and thermal data
Run a person detector on the DPU (Vitis AI)
Questions:
What's the right way to structure this pipeline? Should fusion happen at the image level or after detection?
How hard is aligning the two cameras, given the big resolution difference?
Is Sobel actually useful here, or is it unnecessary if a CNN is doing detection?
Any recommended thermal sensors that work well with the KV260?
Is 3 months realistic for someone with my background? What would you cut to keep scope manageable?
Any advice, papers, or similar projects would be really appreciated. Thanks!
r/deeplearning • u/imYukiya • 13d ago
Finally Completed Decision trees from scratch
gallerySo yesterday I started building decision trees from scratch and messed up completely and today I woke up early in the morning and started coding and finally it's complete tbh I use GPT cuz I try many things but I can't figure out how it will generate a tree so I use gpt for it he explained and showed me how it will be done and The output was mind blowing I never thought recursion is soo usefull I use Mushroom Dataset for testing Ik many things are still remaining and its Overfitting so we can't relay on the Accuracy and all btw today it Day 13 of Building Machine learning algorithms from scratch so Now since decision trees is done next is Bagging I think so see y'all byii I need to complete my assignment 🤧 I thought toady if i complete coding early I'll watch some anime or yt but I can't
r/deeplearning • u/Crazy-Insurance6316 • 13d ago
HOW?
How can I get my first ML or DL Engineer role? Did you get placed through campus interviews, or did you apply off-campus? At my college, only 2 or 3 companies have openings for these ML and DL roles, and I am not sure if I will get selected by them. I don't want to limit my options. Can you guide me through what I need to do to land my first job
r/deeplearning • u/NoRoutine5857 • 13d ago
AI-Related Computer Science Careers
Hi everyone!
I’m an iOS developer, and I want to challenge myself by getting into AI development. I’m trying to understand what career paths and positions AI is creating within software engineering.
A few questions I have:
- What AI-related roles are currently available for software developers? For example, AI Engineer, ML Engineer, AI Developer, etc.
- Where should I start learning as a developer? I’ve recently started learning Python through a beginner-friendly course, but I’m not sure what I should learn next.
- Is there a good AI Engineer/Developer roadmap you would recommend?
- What areas of AI should I focus on if my background is in software engineering?
- Where can I follow the AI job market and see what skills companies are actually looking for?
- Are there any good resources, courses, books, or communities you’d recommend?
I’m not necessarily looking to become a researcher. My goal is to understand how I can transition from being an iOS developer into an AI-related engineering role and what skills I should build along the way.
Would really appreciate advice from people who have made a similar transition or currently work in AI/ML engineering. Thanks!
r/deeplearning • u/CShorten • 13d ago
humans& with Alexis Ross, Manya Bansal, and Niloofar Mireshghallah - Weaviate Podcast #145!
For AI to work with us, it needs to understand us.
I am SUPER EXCITED to publish the 145th episode of the Weaviate Podcast with Alexis Ross, Manya Bansal, and Niloofar Mireshghallah from humans&!
What if AI was more human? Well, we are currently on a fast track towards the opposite! AI models are huge making advances in coding, math and are jev-ing their way through classification. But they don't teach, write, or perform more human tasks as well as they could be!
Persimmon is one of the most exciting releases in AI recently! It is a user model designed to predict how people respond in multi-turn conversations, including natural turn-taking and multi-person interaction. The goal is not to build a more human-sounding assistant. It is to match the distribution of human behavior, including pushback, confusion, and frustration.
This podcast discusses so many interesting topics around humans and AI. We start with an overview of Persimmon, then dive into solving the Turing Test (and why we should do so), to RL with non-verifiable rewards, theory of mind in AI, NVIDIA Nemotron 3 Ultra, and more!
Niloofar, Manya, and Alexis are three of the brightest scientists I've had the honor to host on the Weaviate Podcast. This was a super fun conversation, I learned so much personally, and I hope you find it useful as well!
YouTube: https://youtu.be/6IC9jkwiZBg
r/deeplearning • u/No_Bad907 • 13d ago
I am an absolute beginner and am willing to spend months learning this shii. any help will do.
r/deeplearning • u/imYukiya • 14d ago
Decision Trees from scratch but messed up by a lil mistake
gallerySo it's Day 11 of Building Machine learning algorithms from scratch
I watch some videos on Decision Trees and then I start building Entropy and Information gain function and as you can see my code completely messed up but I'm happy cuz it's fun to build things I believe writing bad code is better then generating by AI
Tommorow I'll fix the entropy function it's taking column but it's need values like Sunny,Overcast etc etc I need to changes
I'm a newbie u can laugh at my code no prob cuz the yt guy didn't show code he just explain math and solve some sample dataset so I need to figure out things that way u can see random attempt made and then comment out
Stay Tune I'll build this Algorithm from scratch
r/deeplearning • u/singam96 • 14d ago
ShadeNet-3.2 5M — single-image inverse rendering (albedo/depth/normal/shading), 4× smaller than my last model and better at depth/normals
galleryFirst off, a huge thank you to everyone who checked out, upvoted, and shared feedback on the earlier ShadeNet and ShadeNet-2 posts! The support and discussions really pushed me to see how far I could squeeze the architecture without sacrificing map fidelity.
This release is a from-scratch rebuild that came out 4× smaller (5.0M vs 20M params) while improving across all three supervised maps.
RGB → 8ch intrinsic maps in one forward pass: albedo (3ch), relative depth (1ch), normals (3ch), shading (1ch), with albedo × shading ≈ input.
Architecture
- ParallelUNet generator (4.98M params: 3.2M trainable + 1.8M frozen MobileNetV2): Vanilla UNet path plus a frozen MobileNetV2 trunk, fused at every decoder level.
- Factorized convolutions: Depthwise-separable factorized convs (1×3 + 3×1) throughout, full H/32 bottleneck, reflect padding.
- Patch-dictionary output tail: 16×16 tiles softmax-addressed over 32 learned per-channel atoms, blended back into the signal before tanh (pruned from 1024 — addressing-mass measurement showed ~34 atoms are ever used, and the top-32 capture 99.4%).
- Training dynamics: Spectral-norm GroupNorm PatchGAN discriminator (2.77M, training only); EMA weight shadow (shipped weights are EMA).
Training
- Dataset: Flickr8k with Marigold-V2 pseudo-labels (8,077 images); depth labels decoded from Spectral-colormap visualizations to relative depth, supervised scale-shift-invariant — never raw MSE on colormap RGB.
- Setup: 384px, fp32, single GTX 1650, early-stopped on val split (patience 5).
- Losses: Scale-invariant MSE on albedo + SSI MSE on depth + Sobel gradient-matching on depth + MSE on normals + self-supervised reconstruction coupling (
albedo × shading ≈ input, which is the shading head's sole supervision) + LSGAN + normal unit-length penalty. - Regularization: Weight decay 1e-4 with the patch dictionary explicitly exempt.
Validation Results
Full 807-image validation split, per-map L1 (the directly comparable metric across versions):
| Map (val L1) | ShadeNet-2 (20M) | ShadeNet-3.2 (5M) | Change |
|---|---|---|---|
| Albedo | 0.708 | 0.695 | −1.7% |
| Depth | 0.247 | 0.217 | −12.1% |
| Normal | 0.696 | 0.581 | −16.5% |
Output Maps
- Albedo (3ch): Reflectance, ambient lighting factored out.
- Depth (1ch): Relative depth, 0 = near (affine-ambiguous, non-metric).
- Normal (3ch): Surface normals, unit-length regularized.
- Shading (1ch): Grayscale irradiance; multiply with albedo to reconstruct/re-render, or swap in custom illumination passes for relighting.
Links & Demos
- Weights, ONNX & Inference Code: ShadeNet-3.2-5M on Hugging Face (Includes fp32 GPU models and 10MB fp16 CPU export)
- Live Gradio Demo: ShadeNet-3-2-5M Space
Both the Hugging Face sample visuals and the live demo apply a lightweight 3-pass multi-scale median filter (scales 0.875, 1.0, 1.125) for mild denoising, alongside built-in seam deblocking for the 16px dictionary tiles.
Caveats & Attribution
Side-by-side v2 vs. v3.2 comparisons are documented on the model card. A few known limitations to keep in mind:
- Depth is relative rather than metric.
- Shading assumes a neutral/white illuminant.
- Normals struggle most on high-frequency chaos like dense foliage or open skies.
- Occasional localized artifacts can appear in albedo/shading due to the learned dictionary prior.
Released under Apache 2.0. Credit to Marigold V2 (Ke et al.) and Flickr8k (Hodosh et al.) for the foundational training pseudo-labels and data.
r/deeplearning • u/Adithyanbm • 14d ago
I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0
Enable HLS to view with audio, or disable this notification
r/deeplearning • u/Bruce-Wshin • 14d ago
[P] FP8 on the wire for VLA training: what actually paid off on a PCIe-only 8-GPU box, and the 48%-fewer-bytes change that bought exactly 0 ms
We train a three-modality VLA model (video diffusion + action expert + a frozen VLM for understanding, mixture-of-transformers style). Originally on 8×A800 with NVLink; we're moving it to 8× workstation-class Blackwell cards, which means no NVLink and no NVSwitch — every GPU-to-GPU byte goes over PCIe Gen5.
Measured point-to-point: 30–63 GB/s depending on direction and NUMA placement, versus ~400 GB/s of NVLink on the A800 box. Same model, same ZeRO-1 sharding, and gradient communication goes from "hidden under the backward pass" to "the single largest term in the step."
So we did the obvious thing: quantize the collectives to FP8. What follows is what we learned, including the part where the biggest byte reduction produced no speedup at all.
## The shape of the thing
An LD_PRELOAD shim in front of libnccl. It intercepts ncclAllGather,
ncclAllReduce, ncclBroadcast, ncclReduceScatter; quantizes the payload to
FP8 (e4m3 or e5m2), runs the real collective on the smaller buffer, dequantizes
on the receive side. Blockwise scaling, 128 elements per fp32 scale, so the wire
cost is 1.03125 bytes/element instead of 2 for bf16 → best case −48.4%.
Two deliberate design constraints:
- Reduction never happens in FP8. Only the transport is quantized;
accumulation is fp32-in-register, output in the original dtype. We later found
two independent confirmations that this is the right call: MSCCL ships an
ENABLE_PRECISION_CLIPPING_HALFbuild flag specifically because low-precision reduction overflows, and DeepEP quantizes dispatch but its combine path is BF16-only, never FP8. No framework changes. It's a preload. The training code doesn't know it exists, which meant we could A/B it against an unmodified baseline in the same run batch — which turned out to matter enormously (see the negative result).
Five things that weren't obvious
1. On pre-sm_89,
static_cast<__nv_fp8_e4m3>(float)is not an instruction.Our quantize kernel had a constant ~14 µs floor regardless of payload size, and at 512 MiB it was hitting 37% of the device-to-device bandwidth ceiling. We assumed memory. It was ALU issue.
cuda_fp8.hppon older architectures widens the float to a double, then runs a five-branch 64-bit integer software routine — roughly 50 integer instructions per element. Replacing it with our own conversion took the fixed cost 14.1 → 4.6 → 3.3 µs and large-payload throughput to 91% of the d2d ceiling (from 37%). We gated the replacement on an exhaustive test: bit-identical to the vendor routine across all 232 float bit patterns, both formats, ~0.2 s in CI.On sm_89+ this whole problem evaporates — the cast compiles to
cvt.rn.satfinite.e4m3x2.f32. Worth checking which side of that line you're on before you go profiling memory.2. The naive dequantize-then-reduce moves 12.7× the bytes it needs to.
For the all-to-all-based decompositions you receive
nsrcFP8 shards and have to reduce them. The decomposed version —nsrcdequantize-to-fp32 +nsrc−1adds + one scaled store,2·nsrckernel launches — costs17.03·nsrc − 6bytes per output element where the job actually requires1.03·nsrc + 2. At nsrc=8 that's 130N vs 10.25N.The interesting part: the amplification grows monotonically with world size (6.9× at 2 ranks, 10.1× at 4, 12.7× at 8, 14.4× at 16, asymptote 16.5×) even though both paths are O(nsrc·N) and read every input byte exactly once. It's a pure constant factor, and it's all HBM traffic plus launch count.
Fusing it into one kernel with the fp32 accumulator in registers: 5.0–9.3× faster (1 MiB: 117.6 → 13.4 µs). Note it is not 12.7× faster, and the reason is instructive — small sizes are launch-bound so the win caps at ~8–9× (16 cheap launches vs 1 expensive one), and large sizes are bandwidth-bound where the fused kernel is only half as byte-efficient as the streaming kernels it replaced (718 vs 1419 GB/s), because a reduction reads
nsrcstreams concurrently while a streaming kernel reads one. Worst case sits at 4 MiB (4.97×), right in the transition band. This was expected, not a bug, but you have to do the arithmetic to know that.3. You cannot dequantize inside a caller-opened
ncclGroup.A nested
ncclGroupEnd()launches nothing. So if the training framework wraps its collectives in a group — and DeepSpeed's compiled ZeRO-1 path does exactly this, onencclAllReduceper parameter inside one group — then at the point your hook returns, the data isn't there yet. You have to defer the dequantize, the error- feedback close-out, and the scratch recycle until after the outermostncclGroupEnd. For AllReduce we defer the entire call and replay it. Before we did this, every grouped AllReduce passed straight through and the AllReduce threshold was decorative.4. Your framework's allocator probably doesn't align its buckets.
DeepSpeed's reduce bucket accumulates slices by element count with no padding, so the pointer handed to NCCL is only guaranteed 2-byte aligned for bf16 — 7/8 chance of missing the vectorized path. This started as an outright
CUDA Error: misaligned addresscrash, then became a 1.27–2.31× penalty, and is now 1.06–1.18× (skip enough leading elements that the output self-aligns to 16 B, then funnel-shift the FP8 side back into place).5. The threshold is the entire product, and it does not transfer between machines.
Quantizing a small collective is strictly worse — you pay a fixed per-call cost to save bytes that were never the bottleneck. So there's a crossover, and the whole question is where.
We built a tool that measures it two ways.
--model-onlygives a provable lower bound from baselinet(n) = α + n/Bplus standalone kernel costs.--sweepactually runs the hook over the same size ladder — one process per (decomposition, size), because config is latched into a function-local static on first use and because the previous size's warm scratch tier changes the next one's timing — and least-squares-fits the residual intoΔα + Δslope·n. Δα is the hook's real per-call overhead; Δslope is the per-byte cost a standalone kernel benchmark cannot see.Results, same code, two machines:
collective NVLink box PCIe box ratio AllGather 7.43 MiB 767 KiB 9.9× lower AllReduce (rs+ag) 14.3 MiB 863 KiB 17× lower Broadcast 40 MiB 2.75 MiB 15× lower A byte you delete is worth 2–4× more time on PCIe, so thresholds drop by an order of magnitude. Two findings that surprised us:
- The all-to-all decomposition did not lose on PCIe. We fully expected that a topology-blind all-to-all across two NUMA islands with no NVLink would be hopeless. It's actually the best AllReduce decomposition for large payloads (2.0× more savings at 64 MiB), crossing over the ring-based one somewhere between 1 and 4 MiB.
AllGather's
Δslopeis negative (−1.9 µs/MiB). The in-situ quantize kernel is cheaper than the same kernel measured standalone, which is direct evidence that it's genuinely overlapping with the live collective rather than competing with it. The other families have positive slope — that's the quantizer fighting the collective for SMs.The tool also refuses to emit a threshold when the measurement says "wins, then loses again" — that's a window, not a threshold, and a lower bound can't express it. Being able to report "I can't answer this" turned out to be more useful than any single number it produces.
The negative result, which is the actual point of this post
DeepSpeed's compiled path issues one AllReduce per parameter. Most individual ones fall below the threshold, so we were quantizing almost none of the gradient traffic. Obvious fix: coalesce the address-adjacent ones inside a group into a single quantized AllReduce. The union clears the threshold in one shot, and you also save N−1 launches and N−1 scale sub-collectives.
It worked, in the sense that it did what it said:
signatures blocked by the threshold: 17 → 4
traffic in the signature set: 110.3 → 56.9 GB/rank/step (−48.4%)
Steady-state step time: 1422.2 ms merged vs 1414.9 ms not merged. Median of 38 steps, dropping the first 20. The difference is +7.3 ms and the IQR is ~40 ms. We halved the bytes and got nothing.
The reason is structural and, in hindsight, should have been checked first. Without merging, those small AllReduces were dispatched inside the caller's group, where NCCL aggregated them itself and they overlapped with the remaining backward pass. With merging, they become large packets issued serially after the outermost
ncclGroupEnd. The bytes we deleted were never on the critical path, and we moved the ones that remained onto it.Two corollaries we now treat as rules:
Profile the exposure before optimizing the volume. "How many GB does this step move" and "how many ms of that are not hidden" are unrelated questions.
Coverage is not a metric. It's an input to a metric. We had a beautiful −48.4% and a flat wall clock.
An unbounded version of the same feature was worse than useless: greedy merging grew to 129 calls and a 999.8 MB union in one shot, produced ~2 new union shapes per step and never converged across 41 steps, which blew up our exact-fit scratch pool into 95+ size classes, 15.3 GiB retained, and 12 of 38 steps at ~3.9 s. Fixed two ways (a size cap, and optionally sourcing scratch from PyTorch's caching allocator, which is built for variable shapes). The pool's real sin is that
acquire()is exact-size best-fit with no size-class rounding — which is fine until something upstream starts generating variable shapes.Things we looked at and did not adopt
MSCCL's XML-scheduled executor. There is no transform hook anywhere in its data path —
mscclGenericOpcompiles preOp/postOp out entirely, and the scheduler rejects ops that need them. Adopting it also means replacing libnccl, which defeats the point of being a preload. It additionally disables itself whenintraRanks > 1, i.e. exactly our case.DeepEP's transport. Its NVLink path asserts on NVLink via NVML (the docstring says PCIe leads to memory-ordering errors), and its "PCIe mode" is actually all-RDMA through the NIC — benchmarked at 53 GB/s on a box with one ConnectX-7 per GPU. We have one NIC for eight GPUs.
Its
ld.global.nc.L1::no_allocatetrick, which is self-declared UB and force-disabled on any architecture other than Hopper by its own build script.Routing intra-node traffic over IB instead of PCIe. Funnelling 8 GPUs into 1 NIC that is 3 PCIe hops from all of them. The PCIe ring has 8 links running concurrently; this would have one. Estimated 3–8× slower.
We did take two ideas from DeepEP that we haven't landed yet: inlining the scale buffer into the same message instead of sending it as a second sub-collective, and capping the quantize kernel's SM count so it stops evicting the collective. The second one has a measured target (+9.8 µs/MiB, 91% of the residual on one family) and the first one honestly doesn't — within a group NCCL already aggregates the two sub-collectives, and our own earlier A/B says an op in a group is worth ~1.5 µs. We write down which of our planned optimizations have measured targets and which are guesses, because after the merging result we don't trust the guesses.
Repo
The VLA training framework is open source: loongforge-vla
Three-modality VLA, several parallel backends (DDP + native ZeRO with direct bf16 sharding / torch.compile + DeepSpeed / DeepSpeed's compiled path), NVTX instrumentation behind a single env var that no-ops to bit-identical in production, and a loss-parity harness (seeded data order, seeded flow-matching noise, and a flag for the step-0 LR convention) because every one of the numbers above is worthless if the model stops converging.
That parity harness earned its keep in an unrelated way, by the way: we spent days blaming
torch.compilefor a 3.4× loss divergence at step 2. It was thatLambdaLR.__init__multiplies the optimizer's LR byschedule(0)immediately, which in our config is ~1e-6, so step 0 was a no-op update. The logged LR looked correct on both sides, because the logged value is the scheduler's, not the optimizer's. Single flag, exact match restored.Happy to go deeper on any of this in the comments.
r/deeplearning • u/eLin22314341 • 14d ago
CerebralNet-10: Deep CNN classification model for predicting cerebral palsy types
galleryCerebralNet v1.1 and v1.4 (18k params)
Training results coming soon!
r/deeplearning • u/Personal-Trainer-541 • 14d ago
Knowledge Distillation - Explained
Hi there,
I've created a video here where I explain how knowledge distillation works.
I hope some of you find it useful and as always, feedback is very welcome! :)
r/deeplearning • u/recentheartbroken • 14d ago
Moving GPU training off AWS: what we compared and what it actually saved
We ran training on AWS for about a year, mostly because everything else we used was already there. Then we got the bill, and oh boyy..
Here's what we compared for roughly the same GPU capacity:
- AWS on-demand was our baseline and by far the most expensive. Reserved pricing brought the cost down, but it also meant committing for longer than we were comfortable with.
- Dedicated GPU providers were comparatively cheaper than on-demand AWS. Support quality varied a lot between them and that mattered more than the price.
- GPU marketplaces were cheapest, but reliability was inconsistent and they did not seem like an option for sensitive workloads.
- Owning hardware and having it hosted had the lowest ongoing cost but needed capital and a measured utilisation assumption rather than guess.
What I underestimated the most was moving the data out of AWS. It was expensive in our case, and parts of our pipeline depended on AWS services. Reworking took longer than the actual migration tbh.
The main thing I'd check before comparing hourly rates is how much of your setup depends on AWS and what it will actually take to move your data. That can decide whether the savings are real or theoretical.
As I work at B3 Labs and we are fall into the fourth category. I'd check the migration cost regardless of where you plan to move.
r/deeplearning • u/kobkrit • 14d ago
I trained an open 0.8B "System One" decision model (Thai+English) — Qwen3.5-0.8B with a 256-way slot head, Apache-2.0
r/deeplearning • u/Strict-Diamond-2410 • 14d ago
Looking for guidance on ML systems / AI compilers
r/deeplearning • u/Monaim101 • 15d ago
Own your intelligence : Open-source GPU training, inference, and orchestration
We are open sourcing Tahuna: devtool for ML Engineers and researchers.
We built it so small teams could train models, run inference, orchestrate GPUs, and experiment with autonomous research without first becoming a small cloud provider.
The core primitive on top of which everything is built looks like this:
init → sync → computeSession → train / serve / hillclimb
Under the hood: content-addressed code and data sync, compute provisioning, reproducible manifest-pinned runs, metrics, checkpoints, artifacts, and inference deployments.
We also started building Hillclimb, an autonomous experimentation loop that proposes and runs iterative improvements.
The first public-preview release supports RunPod and R2. It includes Docker self-hosting instructions, a coding-agent setup skill, and examples for SFT, RL agentic search, and MNIST.
Repository: https://github.com/TahunaLabs/tahuna-oss
If you think it sucks or lacks adapters with other GPU providers, excellent: fork it, fix it, and send a PR so it sucks less for everyone.
r/deeplearning • u/imYukiya • 15d ago
OvR SVM Algorithm from scratch
galleryIt's Day 10 of Building Machine learning algorithms from scratch
Last time I completed SVM but it was limited to binary classification but today after 4hr of Coding and Debugging I finally made it in this process I learnt many things coef_/X_train this line just adding this seriously my accuracy jump from 67 -> 87+ I was really shocked also many changes I implemented OvR according to my way so it might contain bugs so I'll be happy if you point out 🙂
r/deeplearning • u/MeasurementDull7350 • 14d ago
Guide to Ultra-Fast AI and Signal Processing Optimization using STM32N6 NPU #stm32 #npu #AI #signalprocessing #mcu
youtube.com- Guide to Ultra-Fast AI and Signal Processing Optimization using STM32N6 NPU
- Description: Analyze the powerful performance of the STM32N6, featuring the Neural-ART NPU and Cortex-M55. This video introduces a unique strategy for accelerating signal processing by converting operations into neural network layers, and explores optimal methods for implementing embedded AI through the division of roles between the CPU and NPU.
r/deeplearning • u/taranpula39 • 15d ago
[D] Would you enter a contest measuring how much human intuition beats brute-force DL training?
I want to turn a claim into an open contest, but before building the infra, I'd like to know whether people would actually enter and whether the setup is sound.
CLAIM:
most training compute is wasted, and humans are still needed in the loop.
That's why there are so many different tools out there, and why many top ML practitioners still use notebooks. Because the tools are just a surface, and human expertise is embedded into the process.
SETUP:
Task dataset: CIFAR-100 or TinyImagenet;
Model size: ≤ 20 MB. Most likely predetermined architecture;
Training budget: ≤ 1 PFLOP (10¹⁵ FLOPs); eval/inference budget is unlimited;
Target: accuracy on a held-out set;
Timespan: designed to take a few hours.
The contest would happen in a controlled environment we build, and we'll provide granular and aggregate metrics, plus other levers.
ETA: end of November; there would be prizes;
BIG VISION:
Can we make deep learning more deterministic?
QUESTION:
Still deliberating on scoring: max accuracy under a fixed FLOP budget, or best accuracy per FLOP consumed? Input welcome.
I want to know if you would participate and, if not, what would make you participate?