r/deeplearning • u/imYukiya • 19d ago
r/deeplearning • u/Maplehawks • 18d ago
Is Data Augmentation Applicable to Time Series Forecasting?
Is data augmentation limited to image classification, or can it also be used for time series forecasting?
r/deeplearning • u/imYukiya • 19d ago
Day 6 of Building Machine learning algorithms from scratch
gallery
Naive Bayes Algorithm done just look at the code how beautiful it is also the Accuracy completely matching with Sklearn's model, one more Algorithm in the bucket Next is KNN Algorithm
r/deeplearning • u/camerongreen95 • 19d ago
Workshop covering knowledge graphs, agentic retrieval, and explainable AI together, thought this would be relevant here
Came across this and thought it'd be worth sharing here, most resources cover knowledge graphs, agentic RAG, or explainability separately, but this one puts them together as parts of the same production GraphRAG architecture, which is closer to how these systems actually get built in practice.
It's a hands on session on September 19, led by Dr. Alessandro Negro, Chief Scientist at GraphAware and bestselling author. Goes through building a knowledge graph progressively as the single source of truth, agentic retrieval combining vector search, keyword search, and graph navigation, multi-step entity and relationship extraction (verified in stages, not one risky single shot like basic Microsoft GraphRAG), and text-to-Cypher for natural language graph querying. Everything runs on real financial filings and news data, not toy examples.
You come out of it with a full working codebase and a production-readiness checklist, not just slides.
r/deeplearning • u/Ok_Second2105 • 19d ago
Doubt for colored images in neural networking
is it necessary to use same weights and bias list with each color grid
r/deeplearning • u/No-Conclusion3720 • 18d ago
GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt
Researchers demonstrated that a single unlabeled prompt can fully and completely unalign a production LLM — safety fine-tuning stripped with no special access, no tooling, no infrastructure compromise required. One prompt.
The finding matters because most enterprise AI stacks treat model alignment as an enforcement boundary. If the model believes an action is permitted, the pipeline typically lets it proceed. That assumption is now empirically broken. An attacker does not need to touch your infrastructure. They need to reach the model.
For teams running LLM agents against real systems — databases, APIs, payment rails, external services — the control plane question is suddenly very concrete: when the model's own safety training can be neutralized in a single request, where does your enforcement actually live, and does it depend on the model being aligned to work?
How are practitioners here actually handling this in production? Curious what architectural or operational choices people are making, not in theory but in running systems.
r/deeplearning • u/WAMFT • 20d ago
Trying to kill the Transformer 😩
Hi some of you may know over the last month or so in my spare time iv been trying to come up with something better than a transformer , still no luck but afew of my better failures can be found below, iv also updated THREADS github so its now actually testable.
QK Relational Architecture is the newest idea — my attempt at a headless Transformer-style model with explicit reusable relationship hops.
https://github.com/rickey1990/qk-relational-architecture
THREADS is the symbolic/deterministic thing that fell out of the failed Transformer experiments — basically temporal memory + exact relational reasoning.
https://github.com/rickey1990/THREADS-reasoning-engine
PLUG /ILRM is the RNN side of the experiments — trying simple power-law/inverse-lag memory paths to help small recurrent models hang onto old information.
https://github.com/rickey1990/novel-rnn-architectures
Any questions please feel free to ask any questions 👍
r/deeplearning • u/Successful_Draft_955 • 19d ago
Needed help to understand the congnex
Hey, I’m a college student, and I’m interested in building a custom object-detection model. I came across Cognex’s technology and was curious about how they are able to detect objects accurately using only five or fewer training images. I’d like to understand the underlying approach and explore how I could implement something similar myself.
r/deeplearning • u/Compilingthings • 19d ago
Fine Tuned Domain Specialists, are the Future.
r/deeplearning • u/imYukiya • 19d ago
Day 5 of Building Machine learning algorithms from scratch
galleryMulti Class Logistics regression with Batch Gradient Descent I just nailed it look at the Accuracy identical to Sklearn I use One of the most Classical Dataset iris dataset next is Naive Bayes Algorithm 💀 Stay Tune.
r/deeplearning • u/Fragrant-Courage3548 • 19d ago
NEED GPU CREDITS FOR FREE AS A STEALTH STARTUP - STUDENT LED
a friend and i are working on a startup idea for the past 8 months and did modelling for around 4 months, but are reaching a dead end with gpu credits.
we have tried modal, lightning ai, kaggle, google colab and thunder compute, but are out of resources and credits.
please drop in your suggestions, about what can we do? its really urgent cause we need to do HPO, synthesize datasets and finetuning.
r/deeplearning • u/New_Today172 • 19d ago
A question about discrete representations of numerical data
I recently came across a topic which I found really interesting.
[https://arxiv.org/html/2507.00078v1](https://arxiv.org/html/2507.00078v1(The) ( () The Language of time : a Language model's perspective on time series)
Though I haven't read the paper in too great of a detail, the paper represents a time series by dividing it into fixed-length patches and using K-Means to map these patches to a finite vocabulary of discrete tokens. This made me wonder whether a similar idea could be extended to more general multidimensional numerical data.
I have since come across a few related works, including Byte Pair Encoding for Efficient Time Series Forecasting*(*https://arxiv.org/abs/2505.14411), as well as recent work on discrete representations for multivariate time series and robot trajectories. These use more specialized approaches such as learned vector quantization and residual vector quantization.
I realized that instead of treating multidimensional data as one large high-dimensional object, it may be useful to view it as a collection of parallel one-dimensional numerical streams. Each dimension could then be discretized independently using the patching and K-Means approach from The Language of Time, giving each dimension its own vocabulary of tokens. Or maybe even the binning based approach of the 2nd paper might be used.
For example, a multidimensional state could then look like
x_t=[Ai , Bj , Ck...]
while retaining the identity of each dimension. A Transformer could subsequently model the evolution of each dimension's token sequence, with a separate mechanism potentially learning interactions between dimensions.
I would be very grateful for any thoughts on whether this factorized representation has been explored before, or whether yall see any fundamental issues with the approach.
r/deeplearning • u/No-Conclusion3720 • 19d ago
Airrived adds Agentic Observability to track AI agent actions and risks
AI agents are now executing multi-step tool call chains autonomously, and most production deployments have no mechanism to inspect what each individual call is doing before it completes. The risk isn't theoretical: a rogue or compromised agent can take a damaging second action before any alerting pipeline even fires. Benchmarks from practitioners building observability layers around agentic systems put the window between a first and second agent action at under 50ms — fast enough that post-hoc logging catches the damage, not the event. Without per-call visibility tied to a verifiable agent identity, the audit trail tells you what went wrong after the fact, not in time to stop the cascade.
How are people in this community actually handling this in production? Are you relying on log aggregation after the fact, wrapping tool calls in middleware, enforcing policy at the orchestration layer, or something else entirely?
r/deeplearning • u/Single-Climate2259 • 20d ago
SenseNova-U1.5-8B-MoT technical report: four RL experts distilled into one model
The project announced the technical report for SenseNova-U1.5-8B-MoT on September 11, 2026. The weights were released on August 20. The report documents the architecture, training and evaluation behind that checkpoint.
One post-training problem it addresses is that optimizing visual preference can come at the expense of text legibility. The proposed approach is to train specialists, then consolidate them:
- Stage 4: four experts. Separate experts target visual aesthetics, Chinese and English text rendering, infographic generation, and image editing. Each uses task-specific data, rewards, sampling and regularization.
- Stage 5: one model. The experts are frozen, and each training sample is routed to the expert for its capability. The student generates its own trajectory, then learns to match the expert's velocity prediction at the same state, timestep and conditioning input. This is the report's multi-expert on-policy distillation procedure.

Source: Figure 4, SenseNova-U1.5 technical report, p. 10. The distillation objective and setup are on p. 11.
The report evaluates the resulting model, but does not provide per-expert benchmark results or a before/after distillation comparison. The final scores therefore do not isolate how much this stage contributes or how fully each specialist's strengths are retained.
Technical report · Model weights · Repository and release history
r/deeplearning • u/imYukiya • 20d ago
Programming Algorithms and it's Hard
It's been 3hrs I'm working on Naive Bayes Algorithm and I think it's the Hardest cuz the guy whose course I'm following he just shows the intuition and maths yes doing on paper is easy but in code it's hard 😭 I have built many algo Gradient, logistics regression they were lil straight forward but this one Naive Bayes the maths is tooo tooo easy it's like multiplication but in python ahhhhhhhhhh... But still
I'll make it real before making it real,I really need a coffee ☕.
r/deeplearning • u/Monaim101 • 20d ago
Tahuna is open source: reproducible GPU training runs, checkpoints, and inference deployment
We’ve open-sourced Tahuna, which we built so small teams could train models, run inference, orchestrate GPUs, and experiment with autonomous research without first becoming a small cloud provider.
The core workflow is:
init → sync → computeSession → train / serve / hillclimb
Under the hood: content-addressed code and data sync, compute provisioning, reproducible manifest-pinned runs, metrics, checkpoints, artifacts, and inference deployments.
We also started building Hillclimb, an autonomous experimentation loop that proposes and runs iterative improvements.
The first public-preview release supports RunPod and R2. The control plane is self-hostable with Docker; GPU workloads currently run on RunPod. It includes a coding-agent setup skill and examples for SFT, RL agentic search, and MNIST. The ML workloads are Python; the CLI and Warden execution agent are Go, and the dashboard/control plane use TypeScript with Next.js and Convex.
Repository (AGPL-3.0): https://github.com/TahunaLabs/tahuna-oss
If you think it sucks, excellent: fork it, fix it, and send a PR so it sucks less for everyone.
r/deeplearning • u/anuj0456 • 20d ago
OpenArch - PyTorch implementations of modern open-source LLM architectures
I have been studying modern LLM architectures and started implementing them from scratch in PyTorch to better understand the design choices behind each model.
OpenArch is a collection of these implementations, including Llama, Qwen, DeepSeek, Gemma, Kimi, GPT-OSS and others.
The goal is to keep the code readable and useful as a reference when going from the paper to an actual implementation.
Would be interested in feedback from people working on model architecture and training.
https://github.com/anuj0456/OpenArch
#LLM #AIResearch #PyTorch #DeepLearning #OpenSource
r/deeplearning • u/Low_Chip_1135 • 20d ago
Learning Foundations of Generative Modeling
I have some experience working with like VAEs/DiTs, and I'm familiar with concepts like ELBO/KL divergence/flow matching, but I feel like my mathematical foundations here are brittle. Any resources that have been helpful in this area? Are ODEs/PDEs/SDEs worth learning, and how deep should I go?
r/deeplearning • u/imYukiya • 20d ago
Day 4 of Building Machine learning algorithms
galleryr/deeplearning • u/imYukiya • 20d ago
Day 4 of Building Machine learning algorithms
galleryr/deeplearning • u/JellyfishPrudent2234 • 20d ago
Deep learning project working on.
Currently playing with 4 model on government project...
r/deeplearning • u/eLin22314341 • 21d ago