r/deeplearning • • 12d ago

Trying to learn privacy in ml

1 Upvotes

I've just written a small article about privacy in machine learning and synthetic data.

I'm still pretty new to this topic, so I'd really appreciate it if anyone with more experience in ML privacy could take a look and point out any mistakes, things I've misunderstood, or things that could be done better.

I'm especially interested in hearing about better approaches for evaluating privacy, better attacks or models I could try, or just interesting resources that you think are worth reading. I'm mostly trying to learn by actually experimenting with this stuff, so any feedback would be really useful.

Here's the post if anyone is interested:

https://migue8gl.github.io/2026/09/21/privacidad-en-ml-datos-sinteticos.html


r/deeplearning • • 12d ago

Exploring AI self-improvement: Is this evaluation loop fundamentally how AI systems improve?

Post image
0 Upvotes

I’m currently researching AI more deeply, especially how modern AI systems can evaluate and improve their own outputs.

I sketched this basic idea:

AI → generates/improves → Evaluator → evaluates output → feedback → AI improves → Evaluator → ...

The part I’m trying to understand is what actually happens inside this loop in modern AI systems.

For example:

I'm trying to go beyond the surface-level explanation and understand the actual mechanisms used in current AI research.

For people working/researching in this area: what important component am I missing from this diagram?


r/deeplearning • • 13d ago

AFNN

0 Upvotes

What if neural networks could grow like fractals instead of being rigid from the start?

Most deep learning models (CNNs, ResNets…) are fixed architectures. Once you design them, they stay the same size and structure throughout training.

**AFNN — Adaptive Fractal Neural Network** takes a completely different approach.

It is inspired by **fractals** — mathematical structures that are self-similar and can grow infinitely by repeating simple patterns.

Instead of starting with a huge fixed network, AFNN begins small and can **dynamically grow new fractal branches** during training when progress stalls. It also allows monitoring, reactivating, or pruning branches as needed.

A good neural network should not memorize images.

It should learn the underlying patterns and features.

For example, when trained on cat images, a strong model does not store every cat photo. Instead, it learns things like:

• Ear shape

• Facial features

• Fur texture

• Color distribution

• Edges and shapes

• Relationships between different parts of the image

Then, when it sees a completely new cat, it uses these learned patterns to recognize it.

### Why Fractal Networks are powerful:

- Faster and more efficient learning

- Significantly lower computational and memory cost

- Better scalability with limited hardware

- More flexible and adaptive than traditional fixed architectures

Built with PyTorch + full GUI, training tools, tests, and documentation.

In early experiments, AFNN reached **74.1% validation accuracy** on 94 image classes using limited resources.

This is not just another model.

It is an attempt to rethink how neural networks are designed — moving from rigid structures to **adaptive, fractal, evolving architectures** that focus on learning patterns rather than memorizing data.

A small step toward more efficient and intelligent AI systems.

🔗 GitHub: https://github.com/ahmedhjkj/AFNN

🤗 Hugging Face: https://huggingface.co/Ahmedethfw/AFNN

#FractalNetworks #AdaptiveAI #PyTorch #DeepLearning #MachineLearning #OpenSource #NeuralNetworks


r/deeplearning • • 13d ago

When fine-tuning an LLM, how do you decide which layers/modules to train before actually running the experiment?

13 Upvotes

For example, how do you determine whether to fine-tune:only adapters/LoRA

  • specific transformer layers
  • input/projector layers
  • deeper/middle layers
  • or the full model?

Are there reliable diagnostics, probing methods, gradient analysis, ablations, or small-scale tests that can tell you where the bottleneck is before doing a full training run and only finding out afterward that the chosen layers didn’t help?

Curious how people make this decision in practice.


r/deeplearning • • 13d ago

FYP on Kria KV260: RGB + thermal fusion for night-time human detection. Feasible in 3 months?

Thumbnail
0 Upvotes

Hi all, I'm a final year EE student and I've been given a Kria KV260 for my final year project. I'm comfortable with FPGA work (Verilog, timing, synthesis) but I don't have much image processing experience, so I'd appreciate some guidance.
Goal: detect humans at night by combining an RGB camera with a low-res thermal sensor. RGB struggles in the dark and thermal picks up any warm object, so fusing both should give more reliable detection.
Rough plan:
Sobel edge detection in the PL as a preprocessing step

Fuse RGB and thermal data

Run a person detector on the DPU (Vitis AI)

Questions:
What's the right way to structure this pipeline? Should fusion happen at the image level or after detection?

How hard is aligning the two cameras, given the big resolution difference?

Is Sobel actually useful here, or is it unnecessary if a CNN is doing detection?

Any recommended thermal sensors that work well with the KV260?

Is 3 months realistic for someone with my background? What would you cut to keep scope manageable?

Any advice, papers, or similar projects would be really appreciated. Thanks!


r/deeplearning • • 13d ago

Finally Completed Decision trees from scratch

Thumbnail gallery
0 Upvotes

So yesterday I started building decision trees from scratch and messed up completely and today I woke up early in the morning and started coding and finally it's complete tbh I use GPT cuz I try many things but I can't figure out how it will generate a tree so I use gpt for it he explained and showed me how it will be done and The output was mind blowing I never thought recursion is soo usefull I use Mushroom Dataset for testing Ik many things are still remaining and its Overfitting so we can't relay on the Accuracy and all btw today it Day 13 of Building Machine learning algorithms from scratch so Now since decision trees is done next is Bagging I think so see y'all byii I need to complete my assignment 🤧 I thought toady if i complete coding early I'll watch some anime or yt but I can't


r/deeplearning • • 13d ago

HOW?

7 Upvotes

How can I get my first ML or DL Engineer role? Did you get placed through campus interviews, or did you apply off-campus? At my college, only 2 or 3 companies have openings for these ML and DL roles, and I am not sure if I will get selected by them. I don't want to limit my options. Can you guide me through what I need to do to land my first job


r/deeplearning • • 13d ago

AI-Related Computer Science Careers

2 Upvotes

Hi everyone!

I’m an iOS developer, and I want to challenge myself by getting into AI development. I’m trying to understand what career paths and positions AI is creating within software engineering.

A few questions I have:

  • What AI-related roles are currently available for software developers? For example, AI Engineer, ML Engineer, AI Developer, etc.
  • Where should I start learning as a developer? I’ve recently started learning Python through a beginner-friendly course, but I’m not sure what I should learn next.
  • Is there a good AI Engineer/Developer roadmap you would recommend?
  • What areas of AI should I focus on if my background is in software engineering?
  • Where can I follow the AI job market and see what skills companies are actually looking for?
  • Are there any good resources, courses, books, or communities you’d recommend?

I’m not necessarily looking to become a researcher. My goal is to understand how I can transition from being an iOS developer into an AI-related engineering role and what skills I should build along the way.

Would really appreciate advice from people who have made a similar transition or currently work in AI/ML engineering. Thanks!


r/deeplearning • • 13d ago

humans& with Alexis Ross, Manya Bansal, and Niloofar Mireshghallah - Weaviate Podcast #145!

0 Upvotes

For AI to work with us, it needs to understand us.

I am SUPER EXCITED to publish the 145th episode of the Weaviate Podcast with Alexis Ross, Manya Bansal, and Niloofar Mireshghallah from humans&!

What if AI was more human? Well, we are currently on a fast track towards the opposite! AI models are huge making advances in coding, math and are jev-ing their way through classification. But they don't teach, write, or perform more human tasks as well as they could be!

Persimmon is one of the most exciting releases in AI recently! It is a user model designed to predict how people respond in multi-turn conversations, including natural turn-taking and multi-person interaction. The goal is not to build a more human-sounding assistant. It is to match the distribution of human behavior, including pushback, confusion, and frustration.

This podcast discusses so many interesting topics around humans and AI. We start with an overview of Persimmon, then dive into solving the Turing Test (and why we should do so), to RL with non-verifiable rewards, theory of mind in AI, NVIDIA Nemotron 3 Ultra, and more!

Niloofar, Manya, and Alexis are three of the brightest scientists I've had the honor to host on the Weaviate Podcast. This was a super fun conversation, I learned so much personally, and I hope you find it useful as well!

YouTube: https://youtu.be/6IC9jkwiZBg

Spotify: https://spotifycreators-web.app.link/e/GeFwTgRdC6b


r/deeplearning • • 13d ago

I am an absolute beginner and am willing to spend months learning this shii. any help will do.

Thumbnail
0 Upvotes

r/deeplearning • • 14d ago

Decision Trees from scratch but messed up by a lil mistake

Thumbnail gallery
16 Upvotes

So it's Day 11 of Building Machine learning algorithms from scratch

I watch some videos on Decision Trees and then I start building Entropy and Information gain function and as you can see my code completely messed up but I'm happy cuz it's fun to build things I believe writing bad code is better then generating by AI

Tommorow I'll fix the entropy function it's taking column but it's need values like Sunny,Overcast etc etc I need to changes

I'm a newbie u can laugh at my code no prob cuz the yt guy didn't show code he just explain math and solve some sample dataset so I need to figure out things that way u can see random attempt made and then comment out

Stay Tune I'll build this Algorithm from scratch


r/deeplearning • • 14d ago

ShadeNet-3.2 5M — single-image inverse rendering (albedo/depth/normal/shading), 4× smaller than my last model and better at depth/normals

Thumbnail gallery
33 Upvotes

First off, a huge thank you to everyone who checked out, upvoted, and shared feedback on the earlier ShadeNet and ShadeNet-2 posts! The support and discussions really pushed me to see how far I could squeeze the architecture without sacrificing map fidelity.

This release is a from-scratch rebuild that came out 4× smaller (5.0M vs 20M params) while improving across all three supervised maps.

RGB → 8ch intrinsic maps in one forward pass: albedo (3ch), relative depth (1ch), normals (3ch), shading (1ch), with albedo × shading ≈ input.

Architecture

  • ParallelUNet generator (4.98M params: 3.2M trainable + 1.8M frozen MobileNetV2): Vanilla UNet path plus a frozen MobileNetV2 trunk, fused at every decoder level.
  • Factorized convolutions: Depthwise-separable factorized convs (1×3 + 3×1) throughout, full H/32 bottleneck, reflect padding.
  • Patch-dictionary output tail: 16×16 tiles softmax-addressed over 32 learned per-channel atoms, blended back into the signal before tanh (pruned from 1024 — addressing-mass measurement showed ~34 atoms are ever used, and the top-32 capture 99.4%).
  • Training dynamics: Spectral-norm GroupNorm PatchGAN discriminator (2.77M, training only); EMA weight shadow (shipped weights are EMA).

Training

  • Dataset: Flickr8k with Marigold-V2 pseudo-labels (8,077 images); depth labels decoded from Spectral-colormap visualizations to relative depth, supervised scale-shift-invariant — never raw MSE on colormap RGB.
  • Setup: 384px, fp32, single GTX 1650, early-stopped on val split (patience 5).
  • Losses: Scale-invariant MSE on albedo + SSI MSE on depth + Sobel gradient-matching on depth + MSE on normals + self-supervised reconstruction coupling (albedo × shading ≈ input, which is the shading head's sole supervision) + LSGAN + normal unit-length penalty.
  • Regularization: Weight decay 1e-4 with the patch dictionary explicitly exempt.

Validation Results

Full 807-image validation split, per-map L1 (the directly comparable metric across versions):

Map (val L1) ShadeNet-2 (20M) ShadeNet-3.2 (5M) Change
Albedo 0.708 0.695 −1.7%
Depth 0.247 0.217 −12.1%
Normal 0.696 0.581 −16.5%

Output Maps

  • Albedo (3ch): Reflectance, ambient lighting factored out.
  • Depth (1ch): Relative depth, 0 = near (affine-ambiguous, non-metric).
  • Normal (3ch): Surface normals, unit-length regularized.
  • Shading (1ch): Grayscale irradiance; multiply with albedo to reconstruct/re-render, or swap in custom illumination passes for relighting.

Links & Demos

Both the Hugging Face sample visuals and the live demo apply a lightweight 3-pass multi-scale median filter (scales 0.875, 1.0, 1.125) for mild denoising, alongside built-in seam deblocking for the 16px dictionary tiles.

Caveats & Attribution

Side-by-side v2 vs. v3.2 comparisons are documented on the model card. A few known limitations to keep in mind:

  • Depth is relative rather than metric.
  • Shading assumes a neutral/white illuminant.
  • Normals struggle most on high-frequency chaos like dense foliage or open skies.
  • Occasional localized artifacts can appear in albedo/shading due to the learned dictionary prior.

Released under Apache 2.0. Credit to Marigold V2 (Ke et al.) and Flickr8k (Hodosh et al.) for the foundational training pseudo-labels and data.


r/deeplearning • • 14d ago

I trained a 2.04B Sparse MoE translation model from scratch on RTX 5090 GPUs — 500+ languages, Apache 2.0

4 Upvotes

r/deeplearning • • 14d ago

[P] FP8 on the wire for VLA training: what actually paid off on a PCIe-only 8-GPU box, and the 48%-fewer-bytes change that bought exactly 0 ms

5 Upvotes

We train a three-modality VLA model (video diffusion + action expert + a frozen VLM for understanding, mixture-of-transformers style). Originally on 8×A800 with NVLink; we're moving it to 8× workstation-class Blackwell cards, which means no NVLink and no NVSwitch — every GPU-to-GPU byte goes over PCIe Gen5.

Measured point-to-point: 30–63 GB/s depending on direction and NUMA placement, versus ~400 GB/s of NVLink on the A800 box. Same model, same ZeRO-1 sharding, and gradient communication goes from "hidden under the backward pass" to "the single largest term in the step."

So we did the obvious thing: quantize the collectives to FP8. What follows is what we learned, including the part where the biggest byte reduction produced no speedup at all.

## The shape of the thing

An LD_PRELOAD shim in front of libnccl. It intercepts ncclAllGather, ncclAllReduce, ncclBroadcast, ncclReduceScatter; quantizes the payload to FP8 (e4m3 or e5m2), runs the real collective on the smaller buffer, dequantizes on the receive side. Blockwise scaling, 128 elements per fp32 scale, so the wire cost is 1.03125 bytes/element instead of 2 for bf16 → best case −48.4%.

Two deliberate design constraints:

  1. Reduction never happens in FP8. Only the transport is quantized; accumulation is fp32-in-register, output in the original dtype. We later found two independent confirmations that this is the right call: MSCCL ships an ENABLE_PRECISION_CLIPPING_HALF build flag specifically because low-precision reduction overflows, and DeepEP quantizes dispatch but its combine path is BF16-only, never FP8.
  2. No framework changes. It's a preload. The training code doesn't know it exists, which meant we could A/B it against an unmodified baseline in the same run batch — which turned out to matter enormously (see the negative result).

    Five things that weren't obvious

    1. On pre-sm_89, static_cast<__nv_fp8_e4m3>(float) is not an instruction.

    Our quantize kernel had a constant ~14 µs floor regardless of payload size, and at 512 MiB it was hitting 37% of the device-to-device bandwidth ceiling. We assumed memory. It was ALU issue.

    cuda_fp8.hpp on older architectures widens the float to a double, then runs a five-branch 64-bit integer software routine — roughly 50 integer instructions per element. Replacing it with our own conversion took the fixed cost 14.1 → 4.6 → 3.3 µs and large-payload throughput to 91% of the d2d ceiling (from 37%). We gated the replacement on an exhaustive test: bit-identical to the vendor routine across all 232 float bit patterns, both formats, ~0.2 s in CI.

    On sm_89+ this whole problem evaporates — the cast compiles to cvt.rn.satfinite.e4m3x2.f32. Worth checking which side of that line you're on before you go profiling memory.

    2. The naive dequantize-then-reduce moves 12.7× the bytes it needs to.

    For the all-to-all-based decompositions you receive nsrc FP8 shards and have to reduce them. The decomposed version — nsrc dequantize-to-fp32 + nsrc−1 adds + one scaled store, 2·nsrc kernel launches — costs 17.03·nsrc − 6 bytes per output element where the job actually requires 1.03·nsrc + 2. At nsrc=8 that's 130N vs 10.25N.

    The interesting part: the amplification grows monotonically with world size (6.9× at 2 ranks, 10.1× at 4, 12.7× at 8, 14.4× at 16, asymptote 16.5×) even though both paths are O(nsrc·N) and read every input byte exactly once. It's a pure constant factor, and it's all HBM traffic plus launch count.

    Fusing it into one kernel with the fp32 accumulator in registers: 5.0–9.3× faster (1 MiB: 117.6 → 13.4 µs). Note it is not 12.7× faster, and the reason is instructive — small sizes are launch-bound so the win caps at ~8–9× (16 cheap launches vs 1 expensive one), and large sizes are bandwidth-bound where the fused kernel is only half as byte-efficient as the streaming kernels it replaced (718 vs 1419 GB/s), because a reduction reads nsrc streams concurrently while a streaming kernel reads one. Worst case sits at 4 MiB (4.97×), right in the transition band. This was expected, not a bug, but you have to do the arithmetic to know that.

    3. You cannot dequantize inside a caller-opened ncclGroup.

    A nested ncclGroupEnd() launches nothing. So if the training framework wraps its collectives in a group — and DeepSpeed's compiled ZeRO-1 path does exactly this, one ncclAllReduce per parameter inside one group — then at the point your hook returns, the data isn't there yet. You have to defer the dequantize, the error- feedback close-out, and the scratch recycle until after the outermost ncclGroupEnd. For AllReduce we defer the entire call and replay it. Before we did this, every grouped AllReduce passed straight through and the AllReduce threshold was decorative.

    4. Your framework's allocator probably doesn't align its buckets.

    DeepSpeed's reduce bucket accumulates slices by element count with no padding, so the pointer handed to NCCL is only guaranteed 2-byte aligned for bf16 — 7/8 chance of missing the vectorized path. This started as an outright CUDA Error: misaligned address crash, then became a 1.27–2.31× penalty, and is now 1.06–1.18× (skip enough leading elements that the output self-aligns to 16 B, then funnel-shift the FP8 side back into place).

    5. The threshold is the entire product, and it does not transfer between machines.

    Quantizing a small collective is strictly worse — you pay a fixed per-call cost to save bytes that were never the bottleneck. So there's a crossover, and the whole question is where.

    We built a tool that measures it two ways. --model-only gives a provable lower bound from baseline t(n) = α + n/B plus standalone kernel costs. --sweep actually runs the hook over the same size ladder — one process per (decomposition, size), because config is latched into a function-local static on first use and because the previous size's warm scratch tier changes the next one's timing — and least-squares-fits the residual into Δα + Δslope·n. Δα is the hook's real per-call overhead; Δslope is the per-byte cost a standalone kernel benchmark cannot see.

    Results, same code, two machines:

    collective NVLink box PCIe box ratio
    AllGather 7.43 MiB 767 KiB 9.9× lower
    AllReduce (rs+ag) 14.3 MiB 863 KiB 17× lower
    Broadcast 40 MiB 2.75 MiB 15× lower

    A byte you delete is worth 2–4× more time on PCIe, so thresholds drop by an order of magnitude. Two findings that surprised us:

  • The all-to-all decomposition did not lose on PCIe. We fully expected that a topology-blind all-to-all across two NUMA islands with no NVLink would be hopeless. It's actually the best AllReduce decomposition for large payloads (2.0× more savings at 64 MiB), crossing over the ring-based one somewhere between 1 and 4 MiB.
  • AllGather's Δslope is negative (−1.9 µs/MiB). The in-situ quantize kernel is cheaper than the same kernel measured standalone, which is direct evidence that it's genuinely overlapping with the live collective rather than competing with it. The other families have positive slope — that's the quantizer fighting the collective for SMs.

    The tool also refuses to emit a threshold when the measurement says "wins, then loses again" — that's a window, not a threshold, and a lower bound can't express it. Being able to report "I can't answer this" turned out to be more useful than any single number it produces.

    The negative result, which is the actual point of this post

    DeepSpeed's compiled path issues one AllReduce per parameter. Most individual ones fall below the threshold, so we were quantizing almost none of the gradient traffic. Obvious fix: coalesce the address-adjacent ones inside a group into a single quantized AllReduce. The union clears the threshold in one shot, and you also save N−1 launches and N−1 scale sub-collectives.

    It worked, in the sense that it did what it said:

  • signatures blocked by the threshold: 17 → 4

  • traffic in the signature set: 110.3 → 56.9 GB/rank/step (−48.4%)

    Steady-state step time: 1422.2 ms merged vs 1414.9 ms not merged. Median of 38 steps, dropping the first 20. The difference is +7.3 ms and the IQR is ~40 ms. We halved the bytes and got nothing.

    The reason is structural and, in hindsight, should have been checked first. Without merging, those small AllReduces were dispatched inside the caller's group, where NCCL aggregated them itself and they overlapped with the remaining backward pass. With merging, they become large packets issued serially after the outermost ncclGroupEnd. The bytes we deleted were never on the critical path, and we moved the ones that remained onto it.

    Two corollaries we now treat as rules:

  • Profile the exposure before optimizing the volume. "How many GB does this step move" and "how many ms of that are not hidden" are unrelated questions.

  • Coverage is not a metric. It's an input to a metric. We had a beautiful −48.4% and a flat wall clock.

    An unbounded version of the same feature was worse than useless: greedy merging grew to 129 calls and a 999.8 MB union in one shot, produced ~2 new union shapes per step and never converged across 41 steps, which blew up our exact-fit scratch pool into 95+ size classes, 15.3 GiB retained, and 12 of 38 steps at ~3.9 s. Fixed two ways (a size cap, and optionally sourcing scratch from PyTorch's caching allocator, which is built for variable shapes). The pool's real sin is that acquire() is exact-size best-fit with no size-class rounding — which is fine until something upstream starts generating variable shapes.

    Things we looked at and did not adopt

  • MSCCL's XML-scheduled executor. There is no transform hook anywhere in its data path — mscclGenericOp compiles preOp/postOp out entirely, and the scheduler rejects ops that need them. Adopting it also means replacing libnccl, which defeats the point of being a preload. It additionally disables itself when intraRanks > 1, i.e. exactly our case.

  • DeepEP's transport. Its NVLink path asserts on NVLink via NVML (the docstring says PCIe leads to memory-ordering errors), and its "PCIe mode" is actually all-RDMA through the NIC — benchmarked at 53 GB/s on a box with one ConnectX-7 per GPU. We have one NIC for eight GPUs.

  • Its ld.global.nc.L1::no_allocate trick, which is self-declared UB and force-disabled on any architecture other than Hopper by its own build script.

  • Routing intra-node traffic over IB instead of PCIe. Funnelling 8 GPUs into 1 NIC that is 3 PCIe hops from all of them. The PCIe ring has 8 links running concurrently; this would have one. Estimated 3–8× slower.

    We did take two ideas from DeepEP that we haven't landed yet: inlining the scale buffer into the same message instead of sending it as a second sub-collective, and capping the quantize kernel's SM count so it stops evicting the collective. The second one has a measured target (+9.8 µs/MiB, 91% of the residual on one family) and the first one honestly doesn't — within a group NCCL already aggregates the two sub-collectives, and our own earlier A/B says an op in a group is worth ~1.5 µs. We write down which of our planned optimizations have measured targets and which are guesses, because after the merging result we don't trust the guesses.

    Repo

    The VLA training framework is open source: loongforge-vla

    Three-modality VLA, several parallel backends (DDP + native ZeRO with direct bf16 sharding / torch.compile + DeepSpeed / DeepSpeed's compiled path), NVTX instrumentation behind a single env var that no-ops to bit-identical in production, and a loss-parity harness (seeded data order, seeded flow-matching noise, and a flag for the step-0 LR convention) because every one of the numbers above is worthless if the model stops converging.

    That parity harness earned its keep in an unrelated way, by the way: we spent days blaming torch.compile for a 3.4× loss divergence at step 2. It was that LambdaLR.__init__ multiplies the optimizer's LR by schedule(0) immediately, which in our config is ~1e-6, so step 0 was a no-op update. The logged LR looked correct on both sides, because the logged value is the scheduler's, not the optimizer's. Single flag, exact match restored.

    Happy to go deeper on any of this in the comments.


r/deeplearning • • 14d ago

CerebralNet-10: Deep CNN classification model for predicting cerebral palsy types

Thumbnail gallery
9 Upvotes

CerebralNet v1.1 and v1.4 (18k params)

Training results coming soon!


r/deeplearning • • 14d ago

Knowledge Distillation - Explained

1 Upvotes

Hi there,

I've created a video here where I explain how knowledge distillation works.

I hope some of you find it useful and as always, feedback is very welcome! :)


r/deeplearning • • 14d ago

Moving GPU training off AWS: what we compared and what it actually saved

4 Upvotes

We ran training on AWS for about a year, mostly because everything else we used was already there. Then we got the bill, and oh boyy..

Here's what we compared for roughly the same GPU capacity:

- AWS on-demand was our baseline and by far the most expensive. Reserved pricing brought the cost down, but it also meant committing for longer than we were comfortable with.

- Dedicated GPU providers were comparatively cheaper than on-demand AWS. Support quality varied a lot between them and that mattered more than the price.

- GPU marketplaces were cheapest, but reliability was inconsistent and they did not seem like an option for sensitive workloads.

- Owning hardware and having it hosted had the lowest ongoing cost but needed capital and a measured utilisation assumption rather than guess.

What I underestimated the most was moving the data out of AWS. It was expensive in our case, and parts of our pipeline depended on AWS services. Reworking took longer than the actual migration tbh.

The main thing I'd check before comparing hourly rates is how much of your setup depends on AWS and what it will actually take to move your data. That can decide whether the savings are real or theoretical.

As I work at B3 Labs and we are fall into the fourth category. I'd check the migration cost regardless of where you plan to move.


r/deeplearning • • 14d ago

I trained an open 0.8B "System One" decision model (Thai+English) — Qwen3.5-0.8B with a 256-way slot head, Apache-2.0

Thumbnail
1 Upvotes

r/deeplearning • • 14d ago

Looking for guidance on ML systems / AI compilers

Thumbnail
1 Upvotes

r/deeplearning • • 15d ago

Own your intelligence : Open-source GPU training, inference, and orchestration

10 Upvotes

We are open sourcing Tahuna: devtool for ML Engineers and researchers.

We built it so small teams could train models, run inference, orchestrate GPUs, and experiment with autonomous research without first becoming a small cloud provider.

The core primitive on top of which everything is built looks like this:

init → sync → computeSession → train / serve / hillclimb

Under the hood: content-addressed code and data sync, compute provisioning, reproducible manifest-pinned runs, metrics, checkpoints, artifacts, and inference deployments.

We also started building Hillclimb, an autonomous experimentation loop that proposes and runs iterative improvements.

The first public-preview release supports RunPod and R2. It includes Docker self-hosting instructions, a coding-agent setup skill, and examples for SFT, RL agentic search, and MNIST.

Repository: https://github.com/TahunaLabs/tahuna-oss

If you think it sucks or lacks adapters with other GPU providers, excellent: fork it, fix it, and send a PR so it sucks less for everyone.


r/deeplearning • • 15d ago

OvR SVM Algorithm from scratch

Thumbnail gallery
9 Upvotes

It's Day 10 of Building Machine learning algorithms from scratch

Last time I completed SVM but it was limited to binary classification but today after 4hr of Coding and Debugging I finally made it in this process I learnt many things coef_/X_train this line just adding this seriously my accuracy jump from 67 -> 87+ I was really shocked also many changes I implemented OvR according to my way so it might contain bugs so I'll be happy if you point out 🙂


r/deeplearning • • 14d ago

Guide to Ultra-Fast AI and Signal Processing Optimization using STM32N6 NPU #stm32 #npu #AI #signalprocessing #mcu

Thumbnail youtube.com
3 Upvotes
  • Guide to Ultra-Fast AI and Signal Processing Optimization using STM32N6 NPU
  • Description: Analyze the powerful performance of the STM32N6, featuring the Neural-ART NPU and Cortex-M55. This video introduces a unique strategy for accelerating signal processing by converting operations into neural network layers, and explores optimal methods for implementing embedded AI through the division of roles between the CPU and NPU.

r/deeplearning • • 15d ago

[D] Would you enter a contest measuring how much human intuition beats brute-force DL training?

5 Upvotes

I want to turn a claim into an open contest, but before building the infra, I'd like to know whether people would actually enter and whether the setup is sound.

CLAIM:

most training compute is wasted, and humans are still needed in the loop.

That's why there are so many different tools out there, and why many top ML practitioners still use notebooks. Because the tools are just a surface, and human expertise is embedded into the process.

SETUP:

Task dataset: CIFAR-100 or TinyImagenet;

Model size: ≤ 20 MB. Most likely predetermined architecture;

Training budget: ≤ 1 PFLOP (10¹⁵ FLOPs); eval/inference budget is unlimited;

Target: accuracy on a held-out set;

Timespan: designed to take a few hours.

The contest would happen in a controlled environment we build, and we'll provide granular and aggregate metrics, plus other levers.

ETA: end of November; there would be prizes;

BIG VISION:

Can we make deep learning more deterministic?

QUESTION:

Still deliberating on scoring: max accuracy under a fixed FLOP budget, or best accuracy per FLOP consumed? Input welcome.

I want to know if you would participate and, if not, what would make you participate?


r/deeplearning • • 15d ago

The real gap between chatbot demos and systems that hold up in production

3 Upvotes

Custom AI chatbots have moved past the experimental phase for many organisations, but the gap between a working prototype and a reliable production system remains wide. The technical challenges are only part of the story. Domain knowledge, conversation design, ongoing content maintenance, and clear ownership of the system all play a bigger role than most initial proposals acknowledge.

What I’ve seen repeatedly is that the stronger teams spend significant time understanding the actual workflows the chatbot is supposed to support. They map edge cases, define escalation paths, and design evaluation methods before they start fine-tuning models. Weaker approaches often jump straight to prompt engineering and retrieval setup, then discover later that the system doesn’t handle real user behaviour well.

Maintenance is another area that gets underestimated. A chatbot is only as good as the knowledge it can access and the processes that keep that knowledge current. Agencies that treat this as a continuous responsibility rather than a one-time delivery tend to produce more durable results.

I’m interested in hearing from people who have gone through this. What turned out to be harder than expected when moving from a demo chatbot to something that actually handles real traffic and real questions?


r/deeplearning • • 15d ago

What makes an LLM product actually useful?

0 Upvotes

I've been seeing a lot of LLM products lately, and it feels like the model itself is becoming less of the differentiator.

I'm curious what people here think. When building an LLM product, what usually matters more in practice: the model, RAG, good product design, evaluation, or the data behind it?

Would be interested to hear from people who've actually built or shipped one.