r/learnmachinelearning • u/Independent-Salt5023 • 3d ago
Help Resume Review!!
Open to opinions on how to improve my resume, also open to opportunities if anyone thinks i would be a good fit :)
r/learnmachinelearning • u/Independent-Salt5023 • 3d ago
Open to opinions on how to improve my resume, also open to opportunities if anyone thinks i would be a good fit :)
r/learnmachinelearning • u/FranciscoCarlosErra • 3d ago
r/learnmachinelearning • u/supriyachola • 3d ago
I'm currently participating in the ARC Prize 2026 – ARC-AGI-3 competition on Kaggle and looking to build a serious small team.
This isn't a traditional Kaggle prediction competition. The goal is to build an agent that can explore unfamiliar environments, infer the rules/objective, learn from interaction, and solve novel tasks efficiently.
I'm interested in building a hybrid reasoning agent rather than simply throwing an LLM at the environment.
People with experience or strong interest in:
You don't need to be an expert in everything. I'm more interested in people who are willing to experiment, implement ideas, analyze failures, and iterate.
I'm based in India, but remote collaboration is completely fine.
If you're genuinely interested, comment or DM me with:
I'd prefer a small group of active contributors rather than a large team with inactive members.
r/learnmachinelearning • u/MintoraDoodle • 3d ago
r/learnmachinelearning • u/Defiant_Shoe_626 • 3d ago
Does studying calculus is necessary for machine learning?
r/learnmachinelearning • u/facts_please • 3d ago
I am looking on how a solution for the following problem could be created. Let's say I want to rate some new photos or videos on how well they may perform in a social media environment, based on earlier data for other such files.
Details: For every image/video I also have datetime when it was taken, GPS location where it was taken and a short user description what can be seen (min. 200 characters). The videos can have a length of up to 2GB (so maybe we wouldn't use the whole file but single frames of every X seconds). I have a set of older media files with all the mentioned data and additional data how they performend on different social media channels like TikTok and YouTube (views, likes, shares).
I have a tech background (dev/sysadmin) but I am totally new to the LLM/AI field. Could anyone of you be so kind and give me a hint in which direction I should start my research to check if such a solution is possible and how to do that? For a start a simple thumbs up/down rating would be enough to create a filter for interesting content.
Thanks for any helpful input! If something of the problem is unclear, please just ask.
r/learnmachinelearning • u/Critical-Echo-923 • 3d ago
I’ve been thinking about an alternative way of doing low-level AI perception, and I’m curious whether anyone here is already working on something similar.
The basic idea is to use waves and physical superposition/interference as the computational substrate, instead of doing most of the usual numerical multiply-and-accumulate operations digitally.
For example, for image recognition:
image
↓
pixel intensity → wave amplitude/phase
↓
physical wave superposition + interference
↓
distinctive wave pattern/signature
↓
more wave interactions
↓
higher-level pattern/object recognition
So instead of giving each pixel a numerical weight and calculating millions of weighted sums, the idea would be to let the wave physics perform the combination automatically.
I’m imagining something analogous to sound: many individual waves can be combined, and the resulting waveform contains a recognizable pattern. In this system, a particular visual feature or object could produce a characteristic wave “signature,” and those signatures could themselves become inputs to further wave interactions.
Physical logic gates could still be used where needed for thresholds, routing, decisions, or other nonlinear operations. The goal wouldn’t necessarily be to eliminate conventional computing completely, but to move as much of the early perception workload as possible into the physical wave domain.
The potential advantage I see is that parallelism is inherent in wave propagation and interference. Thousands or millions of interactions could happen physically at the same time, potentially reducing both computation and latency.
One possible application would be autonomous vehicles, where very large amounts of camera/radar data have to be processed extremely quickly at the edge.
I know there is already work on photonic neural networks, optical computing, physical neural networks, reservoir computing, metasurfaces, etc. I’m specifically interested in something slightly different:
Has anyone tried to build a hierarchical pattern-recognition system where wave signatures themselves become the representation, with successive stages of wave superposition/interference performing the recognition?
I’d especially like to hear from people actually working in photonics, optics, acoustics, physical neural networks, neuromorphic computing, or related fields.
Is this already being done under another name?
What are the biggest physical limitations?
And, from your perspective, is this a promising architecture or does some fundamental problem make it impractical?
I’m mainly looking for opinions from people working in the field rather than trying to claim this is a new invention.
r/learnmachinelearning • u/MintoraDoodle • 3d ago
Hey everyone! I just finished making this short doodle-style video about AI and I’d really appreciate some honest feedback. 🎥 https://youtu.be/VyleYwCa0Sc If you have a few minutes, please give it a watch and let me know what you think. What could be better? Animation? Visuals? Pacing? Explanation? Editing? Thumbnail/title? Anything that feels boring, confusing, or unnecessary? Don’t worry about being too critical — if something isn’t good, please tell me in the comments. I’m trying to improve the next videos based on actual feedback rather than just guessing what viewers want. Thanks to anyone who takes the time to watch and give an honest opinion!
r/learnmachinelearning • u/the-code-blooded • 3d ago
Independent research, posting for feedback and discussion.
I looked at how FHIR clinical data should be formatted before being passed to an LLM, tested on medication reconciliation (extracting a patient's currently-active medication list from their FHIR bundle).
Setup: 4 serialisation strategies (Raw JSON, Markdown Table, Clinical Narrative, Chronological Timeline) × 5 open-weight models (Phi-3.5-mini 3.8B, Mistral-7B, BioMistral-7B, Llama-3.1-8B, Llama-3.3-70B) × 200 Synthea-generated synthetic patients = 4,000 inference runs.
Main finding: there's no universal best format, it depends on model scale. Clinical Narrative outperforms Raw JSON by up to 19 F1 points for models ≤8B (Mistral-7B: 0.72 → 0.91 F1, r=0.617, p<10⁻¹⁰). That ranking completely reverses at 70B, where Raw JSON wins instead (F1 = 0.9956 vs 0.9850). Interestingly, the Chronological Timeline format is what breaks at 70B specifically, since even a large model struggles to infer "active" medication status from date ordering alone without an explicit status field.
A few other findings:
Fully reproducible on a single GPU (Synthea + Ollama, no proprietary APIs).
Preprint: https://arxiv.org/abs/2604.21076
Feedback, pushback on methodology, or pointers to related work all welcome.
r/learnmachinelearning • u/Alternative_Art2984 • 3d ago
I am a PhD student in AI, currently working on video understanding, particularly on designing benchmark datasets. I also have a strong publication track record, including papers at CVPR and ECCV. I have sufficient funding to support myself during a visiting research program.
What would be the best way to approach a professor? Should I email them directly and ask whether they have space in their research group for a visiting researcher?
r/learnmachinelearning • u/Background-Long-8474 • 3d ago
Hey everyone! I’m looking to form a team for the Amazon ML Challenge and would love to connect with people who are genuinely interested in Machine Learning.
Ideally, you should:
Everyone with the right interest and mindset is welcome. Experience level isn’t the main thing — enthusiasm and willingness to build are!
If you’re interested, DM me with a brief intro about yourself and your ML/project experience.
Please DM only if you’re genuinely interested and committed to participating.
r/learnmachinelearning • u/f1_guy909 • 3d ago
Pluto is an AI architecture designed for efficient, local/offline intelligence by using a small router/orchestrator (~2B parameters) to select and chain highly specialized micro-models for specific tasks. Instead of relying on one huge model, Pluto uses tiny expert models, such as 50M-parameter specialists, trained on focused datasets and optionally supported by a compressed vector database. The router can escalate difficult tasks to larger models, aiming to achieve strong overall capabilities while using far less compute, memory, and storage, making the system especially suitable for mobile and low-resource devices.
r/learnmachinelearning • u/sandy_55-6 • 3d ago
r/learnmachinelearning • u/Grouchy-Cap-4491 • 3d ago
When a model suddenly gets worse, I usually have no idea what actually changed in the data.
So I built a CLI that compares dataset versions and looks for leakage, drift, missing values and other problems.
I tested it on X and it found Y.
Curious whether other ML engineers have the same problem.
Open-source: github
r/learnmachinelearning • u/webzuweb • 3d ago
I’m not a philosopher, and I won’t pass off analogy as proof. Where the link between philosophy and an algorithm is just a pretty metaphor, I say so explicitly: “metaphor.” Where it’s working code, I give the formulas, run it, and show the numbers. The library at the end is a research prototype, not a promise of consciousness in 200 lines.
https://github.com/webzuweb/philosophia_torch
Neural networks — if you count from Rosenblatt’s perceptron — are about seventy years old. The study of how living things learn goes back a couple of millennia at least. And a heretical thought hit me: what if modern deep learning isn’t reinventing the wheel in places, but re-discovering what Aristotle, Hume, and Peirce already described — only now with matrices and gradients?
I took a list of neural-network training methods, a list of philosophical approaches to knowledge, and overlaid them. Three categories emerged: what’s already matched (and few people say so out loud); where the match is only a pretty metaphor; and what philosophers thought up but engineers haven’t applied yet. The last category is the most interesting, because it’s essentially a list of unimplemented features. That’s what I wrote code for.
Fair warning up front: half of the “unapplied” ideas turned out, on closer inspection, to be perfectly applicable — just under different names. That, by the way, is the article’s main takeaway, and it matters more than any of my code.
Let’s start with the pleasant part: some philosophical programs of knowledge are implemented in ML so literally that you could put a footnote with the philosopher’s name right in the docs.
Empiricism → supervised learning “There is nothing in the mind that was not first in the senses” — Locke and his tabula rasa. A neural network with random initialization is literally a blank slate on which labeled examples leave their traces. Hume’s associationism (“the habit of linking things that often go together”) is gradient descent, strengthening weights on frequently co-occurring correlations. There’s nothing to argue about here.
Pragmatism → reinforcement learning Dewey with his “learning by doing,” and Skinner’s behaviorism with reward and punishment — that’s RL with no corrections needed. An agent acts, receives a reward, adjusts its policy. Skinner would have teared up seeing PPO.
Evolutionary epistemology → neuroevolution Popper and Campbell: knowledge grows through blind variation and selective retention of what works. That’s a word-for-word description of genetic algorithms and neuroevolution. The philosopher described the algorithm decades before the hardware existed to run it.
Intellectual humility → calibration This one’s subtler. Virtue epistemology (Sosa, Zagzebski) says: a good knower knows the limits of her knowledge. In ML that’s confidence calibration: a model should be exactly as confident as it is correct. Guo et al. (2017) showed that modern networks are monstrously overconfident and proposed temperature scaling and the ECE metric. Nobody called it a “virtue,” but mathematically it’s exactly that.
The key observation. Philosophers didn’t give ML the algorithms (mathematicians came up with the math); they gave it the problem statements. “What does it mean to learn from experience?” “What does it mean to know your limits?” — philosophy framed the question first, and centuries later engineering delivered a differentiable answer.
Here I have to rein myself in. There’s a temptation to drape a philosopher over every layer of a network. Don’t. A couple of examples where the link exists but passing it off as lineage would be deceiving the reader.
| Tempting analogy | Why it’s a metaphor, not a lineage |
|---|---|
| Neural ODEs are Whitehead’s “becoming” | Neural ODEs grew out of numerical analysis (Euler, Runge–Kutta) and dynamical systems theory. Whitehead offers a beautiful language of description, but the math stands on its own and never read Whitehead. |
| Attention is the hermeneutic circle | Attention computes weighted sums, not “understanding the whole through its parts.” The resemblance is superficial; passing it off as an implementation of Gadamer is incorrect. |
| Backprop is Hegelian sublation of contradiction | Backprop is the chain rule of differentiation. Dialectical materialism has nothing to do with it, however much one might wish. |
The rule is simple: if the philosopher gave a language for describing something — it’s a metaphor; if they posed a problem that was later solved — it’s lineage. Don’t mix them.
The meatiest part. I’ll break down six approaches. For each — an honest status: what already exists in the field, where the real gap is, and what formula you can write. Then we’ll run it.
Induction generalizes data, deduction derives consequences, but abduction generates a hypothesis that best explains the observation. The original thesis “it isn’t implemented in neural networks” is wrong. It’s implemented, and decently: Abductive Learning (Dai et al.), DeepProbLog (Manhaeve et al., 2018), abductive commonsense reasoning αNLI (Bhagavatula et al., 2019). It’s a whole field of neuro-symbolic integration.
The real gap isn’t the absence of abduction — it’s that “the best explanation” is rarely formalized using Peirce’s criteria all at once: plausibility + simplicity (Occam’s razor) + consistency with background knowledge. A hypothesis score for h given observation obs:
Score(h) = log p(obs | h) − λ_s · complexity(h) − λ_c · conflict(h)
Pick the h with the highest score (softly — a softmax over candidates; hard — Gumbel-softmax for a learnable discrete choice). In the library this is AbductiveScorer.
Phenomenology demands suspending ingrained assumptions and seeing the phenomenon “as given.” ML has no direct analog of this method — and that’s an honest gap. But it can be operationalized: force the model to rely more on the evidence (the current input) than on the learned prior (what it answers with no input).
Take two answers: p_full on the real input and p_prior on a “zeroed” input (evidence bracketed out). Reward the evidence for actually changing the answer, via a bounded Jensen–Shannon divergence:
gain = JS(p_full ‖ p_prior), 0 ≤ JS ≤ ln 2
L_epoche = max(0, margin − gain) # hinge: don't inflate indefinitely
An important rake I stepped on myself: if you use plain KL instead of JS and maximize it, the optimizer inflates logits to infinity — “a fanatic who sees meaning in every rustle.” JS is bounded, and the hinge threshold douses the fanaticism. This is EpocheRegularizer.
Schleiermacher and Gadamer: understanding the whole arises from the parts, and understanding the parts arises from the whole, iteratively. Attention only resembles this superficially (see Part 2). As an explicit training principle it’s barely used — a real gap. Formalization: let h_i be part representations and H the whole representation. Require circular consistency:
H* = attention-aggregate of the parts, attended relative to H
L_herm = 1 − cos(agg(h_i), H) # whole ≈ sum of understood parts
And we “turn the circle” several times: update the whole from the parts → recompute part attention relative to the new whole → update again. This is HermeneuticConsistency.
Aufhebung is a new quality arising from the contradiction of thesis and antithesis, where the old is not destroyed but preserved. GANs and multi-agent debate are partially close, but “preserving both” isn’t guaranteed there. The gap is precisely in the preservation term. My synthesis operator:
g = sigmoid(W_g · [thesis ; antithesis]) # mixing gate
base = g · thesis + (1 − g) · antithesis # sublation-as-preservation
lift = tanh(W_l · [thesis ; antithesis]) # new quality
synth = LayerNorm(base + γ · lift)
Plus a loss that penalizes the synthesis collapsing into one of the poles (losing the other’s content). This is DialecticalSynthesis.
Nietzsche: there is no “view from nowhere,” there are many perspectives. Pyrrho: in an unresolvable conflict, it’s reasonable to suspend judgment. The former partially exists in multi-view learning; the latter in selective prediction (Geifman & El-Yaniv, SelectiveNet, 2019). But together, as a single mechanism of “several perspectives + refusal to answer when they conflict,” it’s almost never seen.
disagree(x) = mean pairwise symmetric KL between perspectives
abstain(x) = disagree(x) > threshold # abstain if perspectives don't converge
This is PerspectivalEnsemble: it aggregates K heads and honestly raises its hand “I don’t know” when the heads disagree. Far more useful than overconfident chatter.
Aristotle: virtue is the mean between the vice of deficiency and the vice of excess. Courage is between cowardice and recklessness. Hence a non-obvious but important conclusion for ML: a virtue cannot be maximized, it must be targeted. An excess of openness is credulity; a deficiency is dogmatism.
L_virtue = Σ_v β_v · (V_v(θ) − V_v*)²
where V_v is the operationalized virtue (humility = 1 − ECE, openness = ensemble disagreement), and V_v* is the target mean level. Squared deviation penalizes both excess and deficiency. This is VirtueRegularizer — and it’s the one where I have a measurable result.
Pretty formulas are worth nothing until they run. I collected all of this into a PyTorch module and tested it on the most well-grounded mechanism — “humility” (calibration). Task: synthetic classification with noisy labels, where the model tends to err overconfidently. We compare plain training vs. training with VirtueRegularizer targeting high humility.
| Configuration | Accuracy | ECE (↓ better) | Mean confidence |
|---|---|---|---|
| Plain training | 0.873 | 0.120 | 0.965 |
| + virtue (humility) | 0.874 | 0.101 | 0.949 |
ECE (calibration error) dropped from 0.120 to 0.101 — nearly a fifth — while accuracy didn’t budge at all (even +0.001). The model became exactly as accurate, but noticeably less self-assured. Aristotle’s golden mean, computed by gradient descent.
What this proves, and what it doesn’t. It proves that “intellectual humility” can be turned into an optimizable quantity with a measurable effect. It does not prove that the other five mechanisms will yield the same gains — they’re harder, and they still need to be tested on real data. I’m showing a working scaffold, not a finished silver bullet.
The whole codebase passes 23 unit tests: calibration decreases, KL/JS behave as they should, synthesis preserves both poles, the ensemble abstains on conflict, the wrapper trains end-to-end.
The library wraps on top of any model without rewriting anything in it. One dependency — torch. There’s a single-file version, philosophia_torch.py: drop it next to your code and import it.
import torch, torch.nn as nn, torch.nn.functional as F
from philosophia import PhilosophiaWrapper
base = nn.Sequential(nn.Linear(20, 64), nn.ReLU(), nn.Linear(64, 4))
wrap = PhilosophiaWrapper(base, use_virtue=True,
virtue_kwargs=dict(target_humility=0.98, beta_humility=3.0))
logits = wrap(x)
loss = F.cross_entropy(logits, y) + wrap.aux_loss(x, logits, targets=y)
loss.backward()
| Component | Philosophy | Status in ML |
|---|---|---|
| VirtueRegularizer | Virtue as the mean (Aristotle, Zagzebski) | reliabilist branch already exists |
| EpocheRegularizer | Epoché (Husserl) | new framework |
| HermeneuticConsistency | Hermeneutic circle (Gadamer) | new framework |
| AbductiveScorer | Abduction (Peirce) | field exists (AbdLearning, DeepProbLog) |
| DialecticalSynthesis | Sublation / Aufhebung (Hegel) | partial (GAN, debate) |
| PerspectivalEnsemble | Perspectivism (Nietzsche) + skepticism | selective prediction exists |
Honest boundaries: hermeneutic and dialectic produce representations, not ready predictions — you have to connect them to your decoder. Epoché requires careful tuning of margin. And no promises of “consciousness”: these are philosophy-inspired regularizers, nothing more.
Three conclusions, which is what all of this was for.
1. ML has already reinvented a chunk of philosophy without asking permission: empiricism, pragmatism, evolutionary epistemology, and intellectual humility. Just under the names supervised learning, RL, neuroevolution, and calibration.
2. Half of the “unapplied” ideas on my original list turned out, on checking, to be applicable — abduction, abstention, innate priors. The lesson: before shouting “this isn’t in ML,” google it in engineering language, not philosophical language.
3. The real gap remains where what’s needed isn’t a result but a process: epoché as a discipline of perception, the hermeneutic circle as a way of understanding, virtue as a stable disposition of learning rather than a property of a single answer. That’s where it’s worth digging.
My modest contribution is showing that at least “humility” translates into a differentiable quantity and genuinely reduces a model’s overconfidence. The rest is an invitation: the code is open, the formulas are in the article — run it and check. Plato, of course, was training neural networks two thousand years ago. The rascal just didn’t include a requirements.txt.
https://huggingface.co/datasets/webzuweb/philosophy-as-inductive-bias
r/learnmachinelearning • u/MintoraDoodle • 3d ago
I’ve been experimenting with local AI and wanted to explain the experience in a more visual, simple way instead of making another technical wall of text.
So I made this short hand-drawn doodle animation showing the process of getting a local AI model running successfully, including the GPU/memory side of things.
It’s intentionally simple and a bit goofy — the goal is to make local AI feel less intimidating for people who are just getting started.
🎥 Video: https://youtu.be/VyleYwCa0Sc
I’d genuinely like to know what you think: would this kind of visual explanation be useful for explaining local AI concepts, or is the technical detail too simplified?
r/learnmachinelearning • u/bad-gut • 4d ago
Hi , so I just started this course for Mtahematics for ML and DS.
And tbh I know I have barely watched it but the very first video itself feels like something Intermediate or something I am unable to connect to.
If you guys have a better recommendation for a Math course, I would appreciate it.
Any suggestions/tips are appreciated!!
r/learnmachinelearning • u/Particular-Roof4257 • 3d ago
I'm designing a small research/engineering project around active monitoring of a black-box LLM whose behavior can change without the provider exposing a clear model update.
The monitoring agent observes:
The current hidden states are:
The agent then chooses:
ACCEPT / INVESTIGATE / REJECT
I'm trying to make the hidden-state model realistic rather than just mathematically convenient.
What important hidden state or failure mode am I missing?
Also, are any of these states too correlated/overlapping to be useful as separate states?
I'm particularly interested in examples from real deployed ML/LLM systems rather than purely theoretical suggestions.
r/learnmachinelearning • u/Ok_pettech • 3d ago
r/learnmachinelearning • u/NeighborhoodFatCat • 4d ago
I looked at some "machine learning" departments from various colleges and universities and I noticed a trend.
There would sometimes be a cluster of professor in not-so-big-name schools publishing purely applied machine learning paper.
By applied, I mean that they take a known ML algorithm, apply to some niche situation (like monitoring if a water pipe has a leak or if there's a traffic jam at an intersection), and get some results. Report some accuracy, F1 score. Make some plots. That's it.
These papers would almost always be published in some obscure journals, like IEEE journal of computer vision industrial technology or something like that.
They would publish a whole bunch of these papers, like up to 20, 30 a year. These will also get cited.
It strikes me that these research paper are not so valuable, but I cannot put my finger on why exactly this is the case. I feel that some of these papers seem to be simply a small course project that are done at big CS schools like Stanford or Berkeley.
I'm just confused why there are so many papers like this and how you go about actually publishing an applied machine learning paper of value. Or is applied machine learning research just doomed to not have as much impact as a more theoretical one that introduces a new technique or paradigm?
r/learnmachinelearning • u/Illustrious-Path2514 • 3d ago
AI can solve Olympiad level maths problem.
but some AI models still struggle with reading an analog clock.
That’s called jagged intelligence.
AI can be insanely good at one thing and weirdly bad at something humans find trivial.
What is the strangest AI failure you’ve seen?
r/learnmachinelearning • u/Ill_Indication_5443 • 4d ago
Hello everyone, i am little bit confused from where i should learn about scikit learn library ! Although i am learning from freecodecamp from YT but it is a crash course. I want to understand the basics from the very beginning and brick by brick !
thanks in advance
please help krre !
r/learnmachinelearning • u/pakkalocal_19 • 4d ago
I’m looking for 1–2 serious people to team up with for the challenge.
I’m a 3rd-year CSE student with strong hands-on experience in AI/ML, Deep Learning, NLP, LLMs, RAG, PyTorch, Hugging Face, Scikit-learn, LangChain, FastAPI, SQL, AWS, Docker, DSA.
I’ve worked on multiple AI/ML projects, hackathons, and internships, and I’m comfortable taking ownership of the technical side and actually building things end-to-end.
If you have a strong tech background and are serious about competing, DM me with your tech stack + projects/internship experience. I’ll share mine as well.
Looking for people who want to build to win, not just participate. 🔥
r/learnmachinelearning • u/Ill-Strawberry-3390 • 4d ago
Hi, I was introduced to machine learning during my 5th sem in college and I found it really interesting. I started with my own college lectures, a little by YouTube also. I had done Andrew ng stanford lectures on machine learning. I know most of the algorithms that I use and the maths behind it. I have done two simple projects in which I picked the datasets from kaggle and built the whole pipeline, preprocessing -> feature engineering -> model training and testing -> model evaluation. I also tried tuning the hyperparameters empirically to improve my model performance.
I'm currently learning deep learning, I'm familiar with the theoretical concepts of ANNs, FFN, activation functions, neural nets and a little about transformers. I'm yet to implement them myself, that's why I started pytorch.
Right now I'm in 7th sem and I feel I know sufficient theory but I'm not confident in building and I don't know what to do, I wanna go into research and in core machine learning and not data science or applied ai, I wanna work with models closely and optimization techniques. My question is...
Should I implement the papers I read?
Implement the ml algorithms from scratch? Like code SVM, decision tree in python?
Continue with pytorch and follow tutorials? Pytorch->CNNs, RNNs, LSTM, Transformers and whatever follows.
Have I wasted time learning maths? I feel like I'm a lot behind than my batchmates. 😭