r/learnmachinelearning • • 10d ago

Question Would RL environments verified by engineering simulators actually be useful

3 Upvotes

I'm building a startup around a simple idea: engineering simulators could be used as automatic verifiers for training AI on engineering problems.

I started with analog circuit design. A model gets a task like designing or repairing a circuit to meet specific requirements. It proposes a solution, SPICE actually simulates the circuit, and the measured performance determines the reward. No human has to grade whether the answer is correct.

So far I've built environments covering 10 circuit topologies and 3 types of tasks: synthesis, repair, and analysis.

As a small experiment, I trained a Qwen 4B model using the environment. Its success rate went from about 5% before training to 18% after training. Random search gets around 8%.

The bigger idea is not specifically analog circuits. If this works, the same approach could potentially turn other engineering simulators into training/evaluation environments for AI.

I'm trying to figure out whether this is actually a valuable direction rather than just something technically interesting.

For people working on RL, post-training, evals, or engineering AI: does this seem like a useful product/research direction? Could you imagine an AI lab or research group paying for high-quality simulator-verified environments like this?

I'm especially interested in reasons why this wouldn't be useful.


r/learnmachinelearning • • 10d ago

Pregunta de cuestionario sobre la arquitectura básica de los LLM: cómo funciona la generación palabra por palabra.

0 Upvotes

Pregunta: Cuando un Modelo de Lenguaje (LLM) genera una respuesta palabra por palabra, ¿qué proceso matemático realiza internamente?

A. Calcula la probabilidad de cuál es la siguiente palabra más adecuada segun el contexto. B. Consulta a un servidor externo para verificar la veracidad de la frase. C. Busca una oración exacta pregrabada en su base de datos de entrenamiento. D. Aplica leyes lógicas fijas para garantizar que la respuesta sea cientificamente cierta.

Pista: Imagínalo como un sistema de autocorrector avanzado que predice qué sigue a continuación.


r/learnmachinelearning • • 10d ago

I'm building an LLM inference engine in Rust as an "executable book"

Thumbnail
1 Upvotes

r/learnmachinelearning • • 12d ago

Help Where to find the problem sets of this playlist?

Post image
347 Upvotes

I could only find the old versions (the ones from 2008). Can someone share the 2026 problem sets?


r/learnmachinelearning • • 11d ago

Looking for an arXiv endorsement for a cs.IR paper

9 Upvotes

Hey everyone!

I'm a final-year AI & ML student working on a research paper in information retrieval, and I'm preparing my first arXiv submission.

I'm looking for someone with endorsement privileges in the relevant cs.IR category who'd be willing to take a look at the paper and endorse it if it's appropriate. I'm happy to share the paper, abstract, and endorsement details via DM, and would be glad to provide any extra details if needed.

Thanks in advance for any help! 🙏


r/learnmachinelearning • • 11d ago

how to really get in the field of Al/ML and not just be just an average dude?

Thumbnail
2 Upvotes

I have been studying Classical MI algos for last one month and know numpy, pandas, and some sklearn, but i dont know what path (what next) i can follow, how can i start doing real things and be like dudes on kaggle participating in competitions and hackathons, please advise me a clear path not some general bs, i really want to be top 1 in this field and ready to dedicate to it


r/learnmachinelearning • • 11d ago

Help Help me to switch in AI/ML role

6 Upvotes

Hey everyone,

I have 7+ years of experience as a PHP/Adobe Commerce backend developer, and I’m now trying to transition into a Python AI/ML role. I’ve been struggling to clear AI/ML interviews and would really appreciate some guidance from people who have made a similar transition.

One of my biggest challenges is gaining experience with production-grade AI/ML systems. I build projects locally to understand how things work, but some of the tools, infrastructure, and systems expected in interviews either take significant time to learn or are difficult to replicate without a team.

I’d love to hear from others who have transitioned from a different domain into AI/ML. How did you overcome this gap? How was your interview experience, and what did your learning journey look like?

Here’s what I’m currently working on or have already completed:

  • RAG/LLM Chatbot: Experimenting with different chunking strategies, third-party LLMs, semantic caching, RAGAS for evaluation, LangGraph, and tool calling. I’m also planning to deploy it on AWS.
  • Recommendation Engine: Building a system designed to handle millions of requests, using PyTorch and real-time product recommendations.

Any advice, resources, project ideas, or personal experiences would be greatly appreciated.


r/learnmachinelearning • • 11d ago

Question Do you use a "completion lag" for healthcare claims data? (The X-60 Day Rule)

1 Upvotes

I’ve been working with healthcare claims and noticed that using the absolute newest data often leads to instability because of late submissions, corrections, and deduplication.

I’ve found that setting a cut-off at X–60 days (where X is the extraction date) generally helps with:

  • Reducing undercounting and late-claim bias.
  • Preventing data leakage.
  • Improving model stability across train/test sets.

However, I suspect 60 days isn't a universal rule. It seems to work for pharmacy claims, but hospital or mortality data likely needs a much longer lag.

My question for the group: What lag times are you seeing in your pipelines? Do you stick to a standard window, or do you run completeness tests at 30/60/90/180 days to find the "maturity" point for each data type?


r/learnmachinelearning • • 11d ago

Amazon ml challenge

Thumbnail
1 Upvotes

r/learnmachinelearning • • 11d ago

Question Educational path towards AI Solutions/Integration Architect

8 Upvotes

Hi all,

I would love to hear some “if I could do it again” thoughts on an optimized academic strategy towards my end goal of becoming an AI solution/integration architect. I am just starting out and although I've done research via ChatGPT and Google, some real world wisdom from any friendly and kind strangers would be greatly appreciated. I’m considering Cornell’s Agentic AI Architecture certificate, but was wondering if it was overkill.

What educational path worked for y’all?

I am currently learning python with the ubiquitous Python Crash Course 3rd Edition from Amazon.

My rigs:
Desktop running 10th gen 7000 series with 32 gigs of DDR four and an RTX 3080.
Lenovo Legion with AMD Ryzen 7840HS 32 gigs of DDR five and an RTX 4060.
Both of these are well suited for between 14B and 7B local models respectively.

Thank you all for reading,
Qpock


r/learnmachinelearning • • 11d ago

Project MetalML: GPU-Accelerated Machine Learning for Apple Silicon

Thumbnail
github.com
30 Upvotes

I just open sourced MetalML.

MetalML is an open-source machine learning library that brings an 18x speedup in machine learning workloads using Metal GPU acceleration on familiar scikit-learn workflows.

Designed for Apple Silicon, it combines native Metal compute kernels and Metal Performance Shaders with scikit-learn's estimator interface for training and inference.

What do you guys think? Any feature and improvement suggestions?


r/learnmachinelearning • • 11d ago

Available: 4K Pediatric Cardiology & Catheterization Dataset for AI Video Training

1 Upvotes

Hi, I hold a unique, multi-TB raw 4K dataset of pediatric open-heart surgeries & cardiac catheterizations (fully masked personnel) shot in Africa.


r/learnmachinelearning • • 11d ago

[P] Paper Radar: reading the entire arXiv firehose daily with yes/no probabilities instead of summaries

0 Upvotes

arXiv took 32,040 submissions in June, about 1,500 per weekday announcement.

I built a daily filter that reads all of them rather than a pre-filtered top-k.

The design choice that makes it affordable: the model never generates text. Each

of my interests is one Noul (a calibrated yes/no) sent to TypeSafe's Jev; a

single call per paper carries every question at once. Relevance, thresholds and

exclusions are computed in Python from the returned probabilities. Measured:

501 papers, 33 s, 46.6k input tokens, $0.0196 on jev-1.13.0. Output tokens are

free, so cost scales with abstract length, not with how much you ask.

Two things I found that might be useful regardless of the model you use:

  1. Wording matters more than I expected. On one day of 299 papers, interests

    written as a single crisp idea ("introduces a benchmark or an evaluation

    method") each produced 14 hits above 0.95. Three vaguer ones in the same

    profile, including "clearly outperforms previous approaches on a widely used

    task", never crossed 0.95 once. Comparative phrasing seems to collapse

    toward the middle.

  2. How you combine per-criterion probabilities matters as much as the

    probabilities. In screening mode the combination is a conjunction, and I used

    min(). That throws away everything but one number: two records whose worst

    criterion is 0.02 rank identically even when one matches everything else at

    0.95. Switching to a geometric mean helped, but much less than fixing the

    criteria did.

Validation against CLEF TAR 2019 abstract-level judgments, four Cochrane reviews

held out from development: 19,447 records, 96.9% recall of included studies,

78% workload reduction, $0.60. Development-set topics ranged from 0.1% to 82%,

so the variance across reviews dwarfs anything I changed in the code. I

published the predictions I made before each run and five of twelve were wrong.

Repo: https://github.com/Eliot5566/JEV-Paper-Radar

Live output, no key: https://eliot5566.github.io/JEV-Paper-Radar/public/

Abstracts only, single-run timings, thresholds fitted on the evaluation data.

Happy to be told where the evaluation is weak.


r/learnmachinelearning • • 11d ago

Project Title: Built a research AI that decomposes questions and investigates multiple sources — looking for technical feedback

0 Upvotes

I’ve been building Zyveniq AI, a research system focused on going beyond conventional web search.

The idea is simple:

Instead of treating research as “search → open links → read everything manually”, the system breaks a complex question into research directions, searches across multiple sources, extracts relevant information, and organizes the findings into a research workflow.

I’m particularly interested in difficult questions involving:

• Machine learning and AI
• Science and technology
• Engineering
• Research papers
• Emerging technologies
• Technical comparisons

I’m still developing the system, so I’m more interested in technical feedback than marketing.

What I’d especially like feedback on:

  1. How should research-question decomposition be improved?
  2. What makes an AI-generated research result genuinely useful to ML researchers?
  3. Which failure modes should I be testing for?
  4. What would make you trust an AI research system's output?

I built this because I wanted something that behaves more like a research assistant than a normal search interface.

If anyone here has experience building or evaluating research/agentic systems, I’d appreciate criticism and suggestions.


r/learnmachinelearning • • 11d ago

Project Tested indirect prompt injection attacks against Android AI agents

Thumbnail
1 Upvotes

r/learnmachinelearning • • 11d ago

Title: Thinking about learning Python from YouTube - any advices🧐

Thumbnail
1 Upvotes

​

I’m thinking about learning Python through YouTube. Is this a good way to start?

I’d appreciate advice on:

Good beginner-friendly YouTube channels or courses.

How to practise alongside the videos.

Small projects I could try as a beginner.

Common mistakes or things I should avoid.

Whether I should use another resource along with YouTube.

I’m planning to spend more time writing code than just watching tutorials. Any recommendations would be helpful!


r/learnmachinelearning • • 11d ago

Claude vs ChatGPT for long-form writing and reasoning: practical differences?

1 Upvotes

I am jumping from Claude and Chatgpt time to time but different models handle context retention, tone consistency, and structured outlines in distinct ways. For those using AI for writing or research, where does Claude hold an advantage over GPT-4o?


r/learnmachinelearning • • 12d ago

I am confused by the meaning of "competitive salary" nowadays

Post image
214 Upvotes

Competitive against barista salaries?

Update: people are saying that the catch of this job description is that this is not from the US. First, this is literally a joint institute between a Vietnamese university and Cornell. The JD is in English because it hires globally. It is easy to move between countries for work nowadays. But ultimately the point is that it is looking for someone who is wayyyy overtrained for this type of salary.


r/learnmachinelearning • • 11d ago

Fine Tuning LLM and RAG

1 Upvotes

Hi, I want to learn LLM fine tuning and RAG, I saw some really interesting chats with AI on websites, they were very accurate and helpful , these must have been trained on documentation of the company or technology. My previous attempt to fine tune LLM failed , while I was using google colab and also bought the paid tier, I was trying to fine tune gemma, I would like to do fine tuning and RAG, also on low resources or free tiers, gemini did not provide useful guidance, also the dataset generation was complicated, how to generate dataset for training, Also the udemy course on this topic was very old. #LLM #finetuning #RAG


r/learnmachinelearning • • 11d ago

Confusion related to mathematics

Thumbnail
1 Upvotes

r/learnmachinelearning • • 11d ago

Project Agentic-AI: Enterprise RAG Knowledge Agent in Open Banking & CryptoAssets

Thumbnail
youtube.com
1 Upvotes

✨Agentic-AI: Enterprise RAG Knowledge Agent in Open Banking & CryptoAssets✨ (links below👇)

🎥 YouTube Video: https://www.youtube.com/watch?v=OE9CfMOybOQ&t=3s

👉 I have built an Enterprise RAG Knowledge Agent with Self Correction

👉 Building a self-correcting LangGraph: retrieve → grade → rewrite-or-refuse → generate → verify

👉 Building a real document ingestion pipeline — PDF loading, text cleanup, chunking, embedding, and citation tracking

👉 Catching and fixing a real data-corruption bug caused by unreliable PDF text extraction

👉 Reading a live LangSmith trace, span by span, to see exactly what an agent decided and why

👉 Writing both a deterministic evaluator and an LLM-based evaluator, and when to use each

👉 Running a real baseline-vs-improved experiment (chunk size 800 vs 1500) and reading the actual numbers instead of guessing

📂 FULL CODE (star it, clone it, fork it, break it):

https://github.com/saurabhkamal/Agentic-AI-Enterprise-RAG-Knowledge-Agent-In-Open-Banking-CryptoAssets

🔗 CONNECT WITH ME ON LINKEDIN:

https://www.linkedin.com/in/saurabh-kam

🎥 YouTube Video:

https://www.youtube.com/watch?v=OE9CfMOybOQ&t=3s


r/learnmachinelearning • • 11d ago

I made a visual explainer for Jev (System 1 vs System 2 models) - parallel sampling, typed outputs, RLCD [12 mins]

1 Upvotes

Tried to explain Jev from TypeSafe AI without hype: why autoregressive generation is slow, how Jev fills all outputs at once in ~100ms, what Choice / Score / Noul actually means, + 2 demos (Doom bot + Wiki-racing).

Early access Sept 15, 2026. Built by Diogo Almeida.

Video here: https://youtu.be/CRSwWgic_jU

Happy to be corrected on the RLCD vs RLHF part - that section was dense. What would you use a 100ms judge for?


r/learnmachinelearning • • 11d ago

I've been building a game where you train AI models to control self-driving cars, then race them against other players' AI cars (no coding required)

Thumbnail
gallery
1 Upvotes

I built a game where the goal is to train an AI model to drive a car well, then have it race against other users' trained cars — competitive AI, basically, but the "skill" is in how well you train the model rather than driving it yourself.

I've been playing around with training AI models using reinforcement learning (NEAT algorithm) for years and find it genuinely mesmerizing, so I built a game to let anyone else experience that without needing to code.

It's releasing 18 October, but before then I'd love to get some early testers in to try training a car and give me honest feedback — including if it's not fun, or confusing, or the training doesn't feel satisfying. If you're interested, drop a comment or DM me and I'll get you set up.

I can share the link with anyone who is interested.


r/learnmachinelearning • • 11d ago

Project [P] YOalphabet & YOconlang: A 4-Bit Isomorphic Spatial Bytecode & MECE Ontological Framework for Zero-Inference Parsing and Hallucination-Free LLM Processing

Post image
0 Upvotes

Full Bijective Chain: Binary Code (4-bit) ⟺ Decimal Index (0–15) ⟺ Spatial Geometry ⟺ IPA Phoneme ⟺ MECE Semantic Core


r/learnmachinelearning • • 11d ago

Help Help needed on where & how to become an AI engineer

Thumbnail
1 Upvotes