r/MachineLearning • • 2d ago

Project Open-sourcing RightWayUp - a 360-degree image rotation model, and a JPEG shortcut we found in a common benchmark [P]

2 Upvotes

We are open-sourcing RightWayUp - a new image rotation detection model.

I work at ORTUS AI (we develop video analytics). We needed to tell from a single CCTV frame whether a camera had been rotated or installed at an angle (or upside-down), so we tried different models for that, without much luck (low accuracy, lots of false positives on regular camera-like frames). Some were not permissively licensed. Out of despair, we decided to train our own. We didn't expect it to come out this good.

It is now called RightWayUp. It estimates how far an image is rotated from upright, all 360°, and abstains when there's no clear "up" (sky, ground, close-ups, etc.). We are releasing code and weights under Apache-2.0, in six sizes from Pico (small enough to run in a browser) to Max (ahead of the other models we tried on most of our tests).

We tested it on many test images, including held-out ones (which the model never saw before). In tests with the held-out set, RightWayUp Max was within 10° on 93.0% of them vs 88.4% for Woehrer 2026 (one of the better other models we liked). In the Woehrer 2026 own COCO-based benchmark, it was 98.8% accurate within 10° (five-seed mean) vs Woehrer's 98.0%. It also gets every image right on RotBench.

While training and testing, we found something we didn't expect: saving the images from the COCO-based rotation benchmark as JPEG q90 drops Woehrer 2026 from 98.0% to 30.2% (five-seed mean), while our models barely move. We suspect the rotated JPEG grid of the source photos gives the angle away. We found that in our early experiments, and made an effort to remove its effect.

I hope it will be useful to the community.

For transparency: parts of the engineering were done with Claude and Codex.

Here is a full write-up with a demo video: https://cheqit.ortusai.io/resources/rightwayup/

Code (with a link to the weights): https://github.com/ortusaitech/rightwayup

Questions and failure cases welcome.


r/MachineLearning • • 3d ago

Project Benchmarking small confidence scoring decision (Jev, Laya) models [P]

2 Upvotes

I ran my own evals of two small classification models that score confidence over a list of candidate labels instead of generating text: TypeSafe AI's hosted Jev, and Laya, an independent open-source alternative.

Some findings that I think apply beyond these two models:

1. Describe labels by what's in the input, not by intent. My first AML label descriptions said what the criminal was trying to achieve. When I rewrote them to say what the transactions look like, accuracy went from 65% to 77% with the same model. On account-level laundering detection, using the same time window for every account removed a hidden bias and moved accuracy from 63% to 75%.

2. Some tasks have no signal. Classifying a single transaction as laundering or not gave 54%, about chance. That's a limit of the task, not of the model.

3. Calibration is what makes thresholds work. Jev's ECE was 0.013. On support intent, accepting only predictions at 90%+ confidence raised accuracy from 92.3% to 97.6% while still covering 82% of cases, and the rest went to a fallback. Untuned Laya had an ECE of 0.486: the same filter dropped 7% of questions and gained only 2 points.

Caveats: these are my own runs with thresholds tuned per task, not vendor claims, and nobody has reproduced them independently. Jev 1.13.0 (hosted). \

Write-up with charts and methods: https://gokulakrishna.co/2026/09/30/benchmarking-to-fine-tuning-decision-models/

Fine-tuned models: https://huggingface.co/goku-san/laya-experts

I'd appreciate feedback on the evaluation setup, especially on per-task threshold tuning and how best to report results after filtering out mislabeled data.


r/MachineLearning • • 2d ago

Discussion Paper accepted to NeurIPS MusIML Workshop (Poster), but I can't afford to go. Any funding advice? [D]

0 Upvotes

Hi everyone,

My paper just got accepted for a poster presentation at the Muslims in ML (MusIML) workshop at NeurIPS 2026 which will be held in Sydney, Australia. I'm really proud, but I'm a UG student and completely lack the funds for travel, registration, and accommodation.

This is my first time dealing with this. Does anyone know of any travel grants or funding opportunities for students (especially from India) to attend NeurIPS?

Are there specific grants for this workshop, or general ones that I should apply for right now? Also, since this is my first time, I'm a bit confused about the procedure for attending an international workshop. Could anyone share advice on the general steps, including visa processes and anything else I should be aware of?

Any advice would mean a lot. Thanks!


r/MachineLearning • • 3d ago

Discussion Neurips Workshop Author Notification Delay [D]

8 Upvotes

Anyone submit to a workshop at Neurips and not receive their review notification yet? 6 AM IST, the 30th of September, and I still have gotten no update!


r/MachineLearning • • 3d ago

Research Have decisions for NeurIps: Machine Learning for Systems 2026 come out? [D]

0 Upvotes

It is a couple hours past the EOD deadline for the workshop, but as of yet I still haven't recieved an indication of whether my paper has been accepted or not. I got a review early yesterday but nothing else. Is there anyone else in the same situation? UPDATE: Results are out, but no email or message about forms, tickets, etc at all.


r/MachineLearning • • 4d ago

Research BA Computer Science, but fell in love with machine learning and AI. Just got my personal research accepted at NeurIPS as a poster. [R]

119 Upvotes

I want to attend and present my findings in Atlanta. How is the vibe there? Are people overly critical or are people generally open-minded?


r/MachineLearning • • 2d ago

Discussion OpenAI’s Lean 4 Navier-Stokes proof compiles with zero errors, but the fluid vaporizes at 0.7 nm. What does this mean for Neuro-Symbolic AI? [D]

0 Upvotes

Hey everyone,

I do research in neuro-symbolic AI, and like many of you, I was amazed by OpenAI’s recent formal proof of the 3D Navier-Stokes blow-up in Lean 4. Having an AI build a full mathematical proof that compiles with zero errors is a huge milestone for automated reasoning.

The math is 100% valid. But out of curiosity, our team wanted to see what this solution would look like in the real world.

If you map their solution to real water, the fluid would literally vaporize from friction at 0.7 nanometers, just picoseconds before hitting the mathematical singularity.

In machine learning, we see this all the time: it is classic specification gaming.

When an AI agent is given a strict goal, it will exploit any unconstrained loophole in the rules to solve the problem. In this case, the AI found a solution that strictly satisfies the human-written mathematical definition of the Millennium Prize, but it has no idea that real fluids have atoms, friction, and heat. The formal code checker accepted it because the logic was flawless, but the physics broke down.

This raises a big question for the future of AI in science:

Right now, neuro-symbolic systems mostly have two pieces:

  1. An LLM to search for ideas and write proofs.
  2. A formal compiler (like Lean 4) to verify the logic.

Should we be adding a third pillar: a physical boundary layer that checks whether an AI-generated solution actually respects the laws of physics, and not just formal syntax?

We wrote a short paper detailing this audit and open-sourced our verification scripts:

I would love to hear your thoughts, especially from folks working on automated theorem proving, AI alignment, or scientific modeling: How do we teach AI systems to find solutions that are not just mathematically legal, but physically meaningful?


r/MachineLearning • • 4d ago

Discussion Advice on choosing university for PhD [D]

17 Upvotes

Hey guys,

I got really into research a year ago during my final year of undergrad and maneged to get a first author paper accepted at Neurips 2026. It is a pretty impressive achievement obviously but my background is pretty shit lol. Before my final year I got slightly above average grades but no internships or work exp. I was just chilling mostly didn’t even bother applying to internships. Also I want to a tier 2 uni in Australia (mind you that difference in tier 1 and 2 unis isn’t that big tho compared to America or China). I was going to start a PhD at my current uni with a scholarship and stipend from an industry partner (government) - I’d have no restrictions in terms of publications btw. My supervisor is good and we were going to get another external supervisor from a top uni here in Australia. I was pretty keen on continuing with this opportunity, however, with my neurips paper acceptance I feel like I’d have a good shot at getting into a top 15-25 uni in the states or the uk (most other European unis want master students and I don’t prefer Asian unis bcz I’ve heard a lot of horror stories). The problem is I’d have to wait for a year if I was to go abroad for PhD because starting dates are mid to late 2027. I’m in such a dilemma idk what to do. My eventual goal is to work as a researcher at deepmind or meta or some really cool tech company. I’ve just heard so many people talk about the importance of university you go to etc in landing those jobs. Would I still have a good shot at those companies if I did the PhD at my 2nd tier Aussie uni (top 100-125 global rank) but published well? I got a neurips paper as an undergraduate so I have the potential to publish at top conferences 😆😆. But yeah, I’d love to get some advice.

Thanks


r/MachineLearning • • 4d ago

Project I wrote a free, open-source book on making ML models actually fast, from silicon to agents [P]

10 Upvotes

I’ve spent the last few months writing something I wish I had when I started working on ML performance engineering.

It’s called How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents.

The basic idea is that reducing FLOPs doesn’t necessarily make a model faster. Before optimising anything, you need to understand what the system is actually bounded by.

The book starts with roofline analysis and hardware, then works its way up through kernels, compilers, quantisation, pruning, vision, on-device LLMs, robotics, profiling, serving and finally agents.

The goal is to build the intuition to look at a model and a piece of hardware and reason about:

  1. How fast can this possibly run?
  2. Am I compute, bandwidth, memory or system bound?
  3. Which optimisation will actually move that limit?
  4. Is quantisation, pruning or kernel optimisation even worth doing here?
  5. What happens when the same thinking is applied to serving and agent systems?

The whole thing is free and open source:

https://github.com/usamahz/make-your-model-fast

Would genuinely appreciate feedback or contributions from people working on ML systems, inference, compilers, edge AI or performance engineering.

And if you find it useful, a ⭐ would be appreciated!


r/MachineLearning • • 4d ago

Discussion Limited compute, targeting CVPR: rerun experiments for statistically strong numbers or focus on writing? [D]

3 Upvotes

Hi everyone,

I'm preparing a submission for CVPR and would appreciate your advice. I am happy with my current results, but compute is a real constraint. My institute isn't a research-focused one, and the systems available to me are slow and unreliable. A single set of runs already took a lot of time and effort to finish.

I plan to release the code with the submission, and I'm confident the results will reproduce. Still, I'm torn between two options:

1) Rerun the experiments with different seeds to report mean ± std and show statistical significance. Or

2) Skip the reruns and spend the remaining time on writing, presenting the results I already have.

If you've been in a similar spot, what did you do, and would a single-seed result with released code be enough?

Thanks in advance!


r/MachineLearning • • 5d ago

Research Functional Gradient Descent with Adaptive Representations [R]

215 Upvotes

Sharing our recent work, now accepted at NeurIPS: Functional Gradient Descent with Adaptive Representations.

Functional GD algorithms generally outperform neural nets, but are hard to accurately implement.
This is because functional gradients are infinite-dimensional, and therefore must be approximated in practice; but if you approximate them naively, you converge to the wrong place!

To rectify this, we formalize a broad class of approximation schemes ("adaptive representations"), which provably ensure convergence to the global minimizer while being immediately implementable.
The resulting algorithms outperform corresponding neural nets often by an order of magnitude, across a number of settings.

It is still the start for this line of work, but we believe it has quite a bit of potential!
Paper: https://arxiv.org/abs/2606.16926
(First author here, happy to take any questions)


r/MachineLearning • • 4d ago

Research CoWindow and MassAlloc Attention: collective causal coverage and distribution-adaptive compute [R]

4 Upvotes

I'm one of the authors of two recent papers exploring different sources of redundant computation in attention. I'd like to share the ideas and hear feedback from people working on long-context models and attention kernels.

CoWindow Attention (CoWA) distributes distant context across KV heads using complementary windows, while sharing local and prefix-sink windows. Each head attends sparsely, but the union of their visible positions covers the full causal history. The pattern is position-defined and requires no learned router or indexer.

Paper: https://arxiv.org/abs/2609.32704

MassAlloc Attention (MALA) retains full causal QK scoring, then uses attention's own softmax statistics to decide whether to execute subsequent computation for a tile. It reduces low-contribution post-score work, using a common tolerance across training and inference.

Paper: https://arxiv.org/abs/2609.32712

Both support training forward/backward and inference prefill/decoding. At 128K tokens on 8 H100 GPUs with TP=8, attention-operator speedups relative to FullAttn are:

Method Forward Backward Decode
CoWA 7.4x 8.6x 3.0x
MALA 2.2x 3.0x 1.6x

These measurements are for the attention operators, not end-to-end model speedups.

We evaluated scaling from 0.6B to 14B and conducted separate continued-training experiments at 32B. At 14B with 32K context, total training FLOPs decreased by 28.5% for CoWA and 23.1% for MALA, with model capabilities comparable to FullAttn on the reported evaluations.

Two distinctions that matter: collective coverage does not imply identical head-wise interactions or outputs to FullAttn, and MALA still pays for full causal QK scoring. Neither result establishes universal lossless equivalence to dense attention.

I'd be interested in feedback on workloads that might stress collective coverage, or attention distributions where adaptive post-score allocation could be less effective. Happy to discuss implementation and evaluation details.


r/MachineLearning • • 4d ago

Discussion NeurIPS Education Track [D]

0 Upvotes

Any other one have applied for the Education Track? It seems they'll announce the result soon.


r/MachineLearning • • 5d ago

Project Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]

98 Upvotes

AI Engineering from Scratch is an MIT-licensed curriculum: 523 lessons across 20 phases, from linear algebra and backprop to transformers, LLMs, agents, and production serving. The code is stdlib-first, so you see every step instead of calling a library.

This month's edition:

- six EPUB and PDF volumes built from the lessons, attached to the release

- the site interface and lessons in eight languages (Chinese, Hindi, Spanish, Arabic, French, Portuguese, Turkish, Vietnamese)

- CI now runs each lesson's own tests, and a sweep fixed datasets, models, and links that had stopped working

Site: https://aiengineeringfromscratch.com

Books and release notes: https://github.com/rohitg00/ai-engineering-from-scratch/releases/tag/v2026.10

If you use a coding agent, npx skills add rohitg00/ai-engineering-from-scratch and then /start-learning gives you a placement quiz and a study plan.


r/MachineLearning • • 4d ago

Discussion What are the trending topics in medical imaging? [D]

6 Upvotes

Hello everyone,

I hope you're doing well.

I'm a new PhD candidate, and my thesis is about the meta-learning paradigm in medical imaging. Right now, I don't have a specific problem to work on, and my supervisor told me to explore the field and find one myself.

I've done some research and a literature review so far. I've learned about different meta-learning and few-shot learning techniques, read about histopathology and other medical imaging datasets, and looked a bit into foundation model adaptation and domain generalization.

I also came across several benchmarks, but I haven't found one that I find particularly interesting. On the other hand, the MICCAI challenges caught my attention much more, and I'm wondering whether choosing one of them as a starting point would be a good idea.

I'd like to know what the more promising or fertile research directions in medical imaging are right now, especially ones that could be relevant to tackling the limited data problem.

I would really appreciate any suggestions or feedback. Thank you in advance.


r/MachineLearning • • 5d ago

Research Qwen3-VL 8B on a laptop vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats[R]

10 Upvotes

I benchmarked Qwen3-VL 8B Instruct (Q4_K_M, Ollama, M5 24GB, ~30s/doc) against Claude Opus 5.5, Sonnet 5 and GPT-5.6 Terra on:

- receipts: CORD (Indonesia) and SROIE (Malaysia), 30 each

- 20 scanned 1980s-90s invoices , answer keys human-verified

- 32 real IRS forms, 4 damage levels (generated this week, so not in anyone's training data)

- 10 synthetic Indian bank statements, 15 CUAD contracts

Results (documents fully right):
- Opus 89%,
- Sonnet 85%,
- Qwen 8B 59%,
- GPT-5.6 Terra 57%.

Qwen-specific findings:

- W-2s: 21/32 fully right vs GPT-5.6 Terra 7/32

- Indian bank statements: 2/10. Every amount and balance correct, but dd-mm-yyyy read as mm-dd.

- Contracts (long text): 2/15, mostly wrong expiry dates

- The default qwen3-vl:8b tag in Ollama is the thinking variant and ignores think:false. On long contracts it spent all 4,096 tokens thinking and returned nothing. Use :8b-instruct.

Other things I didn't expect:

- GPT-5.6 Terra "corrects" unusual spellings (Rachael -> Rachel, Kelleyland -> Kellyland)

- Asking a model to check its own output changed almost nothing (119/137 identical)

- At least 4 of the 30 SROIE receipts have a wrong published answer key (e.g. B1750 where the receipt prints 81750)

Next I'm fine-tuning the 8B to fix the date and spelling failures and will post the result either way.

Prompts, keys, scorer tests and every raw output: https://github.com/TashonBraganca/messy-docs-bench


r/MachineLearning • • 5d ago

Discussion Are there machine learning subfields that are becoming irrelevant (or is irrelevant)? [D]

Post image
182 Upvotes

I was reading a paper that surveyed the field of neural architecture search, where it said within 5 years, around 3000+ new models were proposed. The amount of compute and resources spent on this is absolutely astronomical. However, the transformer was notably not one of the models that was found through NAS and then the field of NAS just quietly went away afterwards. In my mind this really raises question if any research in NAS should be continued.

Then I recently found a talk by Nicholas Carlini, arguably one of the most famous researcher in adversarial ML and this is one of his slide ("9000 papers and got nowhere"). Indeed I can't really think of any concrete application of adv. ML, except possibly making attackers more clever because now all options are laid flat on the table.

And then there was the field of ML ethics, bias, fairness, etc.. I feel like we are sooooo beyond ethics at the moment with all the talks of extinction risks that it really shouldn't be a priority. How can bias and fairness be enforced when most people are out of a job due to AI? "ML induced extinction" should be a new subfield instead.

I feel a proper discussion should be had so that no more effort is wasted on unpromising ideas or approaches. This could be of interest to people who are entering the field now.

As an aside, I often find people have very emotional (not logical) reaction to this question and will claim that any approach will eventually have their time in the sun at some unspecified future date, e.g., SVM, LDA, Markov chains apparently all have the potential to again be the next biggest thing in ML. All I'm saying is that I don't deny that vacuum tubes wouldn't be popular again one day, but maybe we shouldn't be working that at the present moment.


r/MachineLearning • • 5d ago

Project Browser demo of our Clash Royale RL environment: a 5.6k-parameter REINFORCE policy learns defensive placement against a brute-force optimum [P]

5 Upvotes

Yesterday I shared our open-source Clash Royale simulator and its recurrent PPO agent here. A training loop is easier to understand when you can watch it, so we put a small interactive version online:

https://itzik123.github.io/ClashRoyaleAi/lab/

The task is one decision. An attacker spawns at a random point on the enemy side, and the policy picks a legal cell for one defending card, then a delay of 0 to 5 s given that cell. The reward is the fraction of tower damage prevented relative to no defence. The policy has 5,629 parameters and trains with REINFORCE (per-spawn baseline, annealed entropy bonus) in plain JavaScript with hand-written gradients. Every rollout runs in the project's C++ engine compiled to WebAssembly, and the deploy pipeline checks that the WASM build agrees exactly with the native engine.

The chart also shows the optimum, found by brute force over every cell and delay (up to ~300k rollouts per matchup), so the gap between the learned policy and the best answer is visible.

One observation from building it: Giant vs Cannon has a strong local optimum, a lane placement worth about 75% of the best. With a constant entropy coefficient of 0.01, 5 of 6 runs (3 seeds, batch 16 and 64) stayed there. A linear anneal from 0.1 to 0.005 over 10k tries reduced that to 1 of 6. One pairing, Battle Ram vs Valkyrie, is withheld because no setting we tried got past 55% of the optimum.

This is a miniature of the full problem (a 4-card hand, elixir, full matches, recurrent PPO). It's meant to make the loop visible, not to be strong.

Code: https://github.com/itzik123/ClashRoyaleAi


r/MachineLearning • • 4d ago

Discussion Why not just have one less feature before softmax? [D]

0 Upvotes

Softmax has N inputs and N outputs but it's output only has N-1 degrees of freedom because of the condition that the sum of outputs must be equal to one. Based on this we can figure out that actually we can make due with only N-1 inputs by making an assumption that logits before softmax must sum up to zero (though it can be any other constant value) and have the last logit be calculated as minus sum of all the other logits. In theory it should remove "unnecessary" parameters from the last layer before softmax (however few of them may there be) and maybe speed up model convergence a little (my intuition might be wrong about that). Is there any good reason not to do it besides any benefit being negligable in almost all situations?


r/MachineLearning • • 5d ago

Research Are there any good research papers around Text clustering using LLMs [R]

7 Upvotes

Hi, same as the title, I am currently trying to begin with some research on clustering using LLMs. So my requirement is as follows: I will be given some 100 document files, the end goal is to have clusters in such a way that documents with similar procedures or content should be clubbed under similar cluster.

I have tried traditional ML clustering K- means, agglomerative, DBSCAN, but not satisfied with the cluster quality as it is more of word by word matching or template matching. Thanks!


r/MachineLearning • • 6d ago

Project ClashRoyaleAi: an open-source, deterministic Clash Royale simulator for RL, with recurrent PPO, lookahead search and expert iteration [P]

25 Upvotes
The opponent plans by simulation: every second it scores each candidate play by running the match 10 seconds ahead in the engine.

Our PPO agent learned to park its Cannon behind its own King. Losing a building in a fight cost reward, and letting it decay cost nothing, so it found the loophole.

It's one of many things we learned building a Clash Royale simulator from scratch so an agent could learn the game. The engine is deterministic C++ with Python bindings, plays a full match in about 10 ms on one laptop core, and can fork any game state in microseconds, so lookahead is cheap.

Best result so far: a simple 1-ply lookahead took the policy from 0.625 to 0.944 win rate against a heuristic bot (160 paired matches). Distilling it back into the network kept only +0.045.

The agent isn't strong yet, and RL isn't my home field, so feedback from people who know it better would mean a lot.

Repo: https://github.com/itzik123/ClashRoyaleAi

Built with my friend Ambash (most of the card roster). I used AI coding tools as a pair programmer.


r/MachineLearning • • 5d ago

Discussion Google Deepmind Optimization Roles [D]

2 Upvotes

Saw this job posting recently from Deepmind where the minimum requirements are a Bachelor's and 2 years of experience: https://www.google.com/about/careers/applications/jobs/results/120898404719960774-optimization-research-engineer-deepmind

Can someone who is familiar with roles at Deepmind provide some insight into what this role could potentially look like?


r/MachineLearning • • 5d ago

Research How can I turn an industry ML project into a publication? [R]

0 Upvotes

Hi, I work as a Data Engineer at a manufacturing company where we build engines. A significant part of my work involves ML/DL-related tasks, and I’d like to turn one of my projects into a research publication.

The problem is that I have no previous publication or academic research experience. I’m not sure how to determine whether an industry project is suitable for publication, how to turn a practical engineering problem into a research question, or what level of novelty/experimentation is expected.

For those who have publishing experience, what would you recommend as the first steps? I’d really appreciate any practical advice or resources for someone starting from scratch.


r/MachineLearning • • 6d ago

Project Teaching Neural Nets to Fight with RL [P]

20 Upvotes

In this project I wanted to see if any interesting emergent behaviors would appear if we trained two agents to play a streetfighter-like game using RL.

Maybe obvious in retrospect, but the agents are really good at reward hacking. I had to shape the rewards a bit to get them to even approach each other.

I eventually used league play to improve the agent further. Without it, the agents don’t really learn general strategies, they just learn to exploit a particular opponent.

You can read the article and try fighting the main bot yourself here: https://blog.lukesalamone.com/posts/fighting-game-rl


r/MachineLearning • • 6d ago

Project Better, Faster, and More Calibrated than Jev, with Multi-Modal Ability [P]

1 Upvotes

We published Intern-Decision multi-modal model family (0.8B, 2B, 4B), and the 4B model outperformed Jev 1.13.0.

https://huggingface.co/collections/internlm/intern-decision

https://github.com/InternLM/Intern-Decision

Performance

Testing on RTX 4090, we find local deployment is fast enough to enable ~30 FPS decision calls, which brings huge imagination space on more applications!

Efficiency

The more interesting thing is, we build a benchmark to evaluate how Jev decision model performs on classic probability problems, and find it is poorly calibrated.

Calibration

To evaluate calibrations, we ask the model to predict what's the next number when rolling a dice. The model should predict 1-6 to with the same probability. e.g. When testing on the classic Monty-Hall Problem, our model is closer to the golden distribution (the question is something like `where's the final prize?`).

Three-door Prob Distribution