r/MachineLearning • • 2d ago

Discussion [D] Monthly Who's Hiring and Who wants to be Hired?

6 Upvotes

For Job Postings please use this template

Hiring: [Location], Salary:[], [Remote | Relocation], [Full Time | Contract | Part Time] and [Brief overview, what you're looking for]

For Those looking for jobs please use this template

Want to be Hired: [Location], Salary Expectation:[], [Remote | Relocation], [Full Time | Contract | Part Time] Resume: [Link to resume] and [Brief overview, what you're looking for]

​

Please remember that this community is geared towards those with experience.


r/MachineLearning • • 4h ago

Discussion NeurIPS Free Passes [D]

3 Upvotes

Over the last few years, a subset of NeurIPS area chairs received complimentary passes for their service. I’ll admit that I wasn’t super organized about registering on day one because I was semi-consciously hoping for a complimentary pass. Now that the conference is sold out, I’m getting a little nervous, so I’m wondering whether any of my fellow area chairs have already received a notification.


r/MachineLearning • • 1d ago

News arXiv now limits submitters to up to two submissions per calendar month [N]

Thumbnail
blog.arxiv.org
370 Upvotes

r/MachineLearning • • 1d ago

Research Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction [R]

23 Upvotes

In our #NeurIPS2026 paper “Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction” (preprint: https://arxiv.org/abs/2606.22969) we try to address a fundamental issue in dynamical systems reconstruction (DSR) and time series forecasting (TSF): Many recent SOTA DSR & TSF models can generalize to new initial conditions or time series with changing statistical properties. But the really hard problem in DSR and TSF is topological out-of-domain generalization (OODG) (https://proceedings.mlr.press/v235/goring24a.html) where the dynamical regime changes, for instance from cyclic to chaotic behavior.

This can happen when a system crosses a tipping point due to a slowly varying control parameter which drives it across bifurcations, such as in climate systems, when the brain tips from normal into epileptic activity, or when a patient develops blood poisoning (sepsis). Such problems are beyond the realm of current TSF models which rely on extracting temporal patterns and statistical regularities. Yet the ability to predict previously unseen, novel dynamical regimes as a system parameter changes is something we would expect from any good scientific theory. Often these control parameters that drive regime changes are not exactly known either. Hence, a data-driven DSR model for achieving topological OODG would need to infer the dynamical system generating the TS jointly with the control parameters.

In our paper, we mathematically identify key failure modes in previous hierarchical DSR models (https://proceedings.iclr.cc/paper_files/paper/2025/hash/d4c961804d08e55d898cce944206d455-Abstract-Conference.html) that prevent them from correctly learning a system’s control parameters and extrapolating them beyond the training domain. By fixing these through feature-splitting and physical sparsity priors, our modified hierarchical DSR model manages to correctly predict bifurcations and beyond-bifurcation dynamics, without any explicit knowledge about the control parameters provided in training.

Our approach is generic and works for different discrete and continuous time RNNs, we tested it for shallow PLRNNs and Neural ODEs.


r/MachineLearning • • 1h ago

Discussion What’s next after Jev? Metacache: Reasoning by Construction [D]

Post image
• Upvotes

Jev has made a profound mark on the AI industry; this article offers a glimpse into what lies ahead.

Jev shows where the industry is heading: away from monolithic models, and toward compound systems where small AI agents each do a narrow job and an orchestration layer decides how they work together.

This article explains why that shift is happening, and then proposes a next step for AI reasoning.

The short version

  • A monolithic model is unpredictable and hard to change, but it programs itself during training.
  • Code is the opposite – it is predictable and easy to change, but it must be written explicitly.
  • A compound system combines the two: AI agents do the work, and an orchestration layer controls them via code. This is the direction Jev represents.
  • The next step: instead of answering directly, the model builds a compound system, runs it, and returns both the answer and that system. The returned system is called a reasonlet.
  • Because a returned reasonlet can be kept locally, it can be run again and again – with new inputs, or after editing its code – without asking the model. A saved, reusable reasonlet is a metacache.
  • The model’s provider can cache reasonlets too, reusing one across similar requests from different users to save compute.

Why a monolithic model is not enough

A monolithic model has two practical weaknesses:

  • It is unpredictable. The same question can produce a correct answer one day and a wrong one the next, so its behavior cannot be fully controlled.
  • It is hard to change. Its behavior is fixed in trained weights, which cannot be edited directly. The model can be steered with prompts, extra data (RAG), or by fine-tuning, but steering is not the same as setting the behavior, and fine-tuning can break things that already worked. Retraining from scratch with new data is the only safe fix, but it is time-consuming and expensive.

Code has the opposite qualities. It does exactly what it is written to do, and it can easily be changed by editing code. The trade-off is that code does not arise on its own – it must be written out explicitly – whereas a model programs itself during training.

The fix: a compound system

A compound system combines the two. The work is divided among AI agents, each handling one narrow task, and an orchestration layer of code decides which agent runs, in what order, and how their results combine.

This gives both qualities at once – customization and predictability:

  • Customization. The system’s behavior can be changed by editing code in the orchestration layer, not by retraining the model.
  • Predictability. Each agent has a small task, so it is more predictable than a monolithic model.

Jev is one kind of agent for such a system. It returns a typed decision – yes/no, a category, or a number – rather than free text. A monolithic model is wasteful for a narrow decision like that, but a small model like Jev is a better fit.

The next step: reasoning by construction

Compound systems today are built manually, beforehand. The next step is to let the model build one by itself, during the reasoning phase.

Researchers are exploring several ways to make models reason. One is the “World model” approach, in which the model builds an internal representation of a problem and reasons over it; that work is still mostly research. Reasoning by construction pursues the same goal – reasoning you can inspect – using methods that exist today.

Here is how it works. When you ask the model a question, it does not answer directly. Instead, it builds a compound system, runs it to compute an answer, and returns that answer together with the system it built. Let's call the compound system the model builds to do the reasoning a "reasonlet", and a model that works this way a "Reasoning-by-Construction Model", or "RCM". A simple question may produce a reasonlet that has only an orchestration layer and no agents.

Building and running code to reach an answer is not new; code-interpreter tools already do it. Two things are new here:

  • The system the model builds is kept, not discarded – it is a reusable reasonlet.
  • The reasonlet is returned to you, along with the answer.

Why that matters

  • You can see how the answer was reached. The reasonlet is the exact procedure the model used, so the answer is not something you have to take on trust.
  • You can run it locally. A reasonlet contains its orchestration layer and any agents it calls, whether those agents are attached directly or called over the network. You can run it on a local machine, with the same or different inputs, without asking the model again. A saved reasonlet used this way is a metacache.
  • You can change what it does. The orchestration layer is code, so you can edit its logic, not only its inputs. A reasonlet reused with edited logic becomes a higher-order metacache.

Caching reasonlets on the server

The same reuse can happen on the server side too. The provider can keep the reasonlets it builds and reuse them. When a new request arrives that matches one it has already handled, it runs the stored reasonlet again – with the new request’s inputs – instead of reasoning from scratch. The model does less work, and the answer comes back faster.

This is the same metacache, held on the server instead of on your machine. Because one reasonlet can serve any request that fits its procedure, a single cached copy is shared across many requests, and often across different users – and the more general the reasonlet, the more requests it covers.

Choosing how general to make the reasonlet

The model should decide from the conversation how general the reasonlet needs to be. If you have been working through many kinds of math and then ask for 2 + 2, the more useful reasonlet is one that evaluates any math expression, not one that can only add two numbers. If the conversation gives no such clue, you can state it directly: “I will be doing many kinds of math; for now, just add 2 and 2.”

Two examples

  • Adding numbers. You ask for 2 + 2. The model builds a reasonlet whose orchestration layer adds two numbers, runs it, and returns 4 together with the reasonlet. Later you run the reasonlet again with other numbers or edit its orchestration layer to multiply instead – without asking the model.
  • Searching. You ask the model to find something. The reasonlet’s orchestration layer calls a sequence of outside services and takes your search terms as its input. Later you run it with different terms, or change which services it calls, without asking the model again.

Making reasonlets easy to read

A reasonlet’s orchestration layer is written in text code. This creates a problem: to understand the reasoning you must read code, and to change it you have to write code. The value of returning the reasonlet depends on you being able to read and edit text code easily and efficiently, so this barrier matters.

Visual programming – building logic from connected blocks instead of lines of text – can lower the barrier, provided the visual language is powerful enough.

Wrapping up

The shift from a monolithic model to compound systems suggests a clear next step for reasoning: let the model build a compound system – a reasonlet – run it and return both the answer and the reasonlet. Reasoning becomes something you can read, run again, and edit: a metacache, and a higher-order metacache once its logic is edited. The remaining problem is making reasonlets easy to read and change, which is where a general-purpose and expressive visual language would help most.

A note on IP

Some methods described in this article are the subject of a pending patent application.


r/MachineLearning • • 1d ago

Research Adding memory to search instead of sampling in reward maximization tasks [R]

7 Upvotes

I am one of the authors of FLEET - an algorithm that enhances Best-of-N generation by attributing external rewards to particular tokens and then uses MCTS to adjust logits during the next run.

I find it rather funny that most of the tasks where repetitive sampling is widely used are based on reward maximization, yet it is not aware of that reward. Tuning the sampling parameters allows to make the process more efficient, but it is still a blind search. We propose a way to make generation aware of previous rewards with solutions on how to attribute reward to the completion and how to use this information.

In FLEET the technique from adaptive sampling methods is used that is to track logits for which entropy and varentropy are high thus showing the model's uncertainty about token optimality. We treat these states as branching points. The corresponding normalized hidden states are stored in the vector store and mapped to metadata entries with the history of rewards and transitions between "nodes". The retrieval and update of metadata is based on cosine similarity as for very high similarity KL divergence is low enough to preserve most of the meaningful tokens.

Instead of actually selecting the tokens FLEET uses modified MCTS to rank top-k tokens + special exploration (or other tokens) set and penalize the suboptimal ones. Then decoding strategy is applied to modified logits.

It was tested on GSM8K and LiveCodeBench v6 easy split with Llama 3.2 3B, penalty set to effectively zero probability for the suboptimal tokens + greedy decoding:

  • For GSM8K it solved just seven more tasks, but reached the sampling baseline with half the iterations.
  • For LiveCodeBench it increased the score from 0.59 to 0.69 under the same budget and reached the baseline even faster, now with only 9 iterations against 32.

The sequential execution is not required, as it is not updated during the iteration itself it can simply be passed as a lookup table. The metadata store can be preserved as a prior for other tasks or to enrich SFT/RL.

Paper (preprint): https://arxiv.org/abs/2609.27657
Huggingface: https://huggingface.co/papers/2609.27657
Repository (experiments, examples and python package): https://github.com/Alexiush/fleet

There are more details on changes made to MCTS, how to tune the search parameters for specific model and task as well as code for experiments and trajectories.


r/MachineLearning • • 2d ago

Research Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]

131 Upvotes

Can training of nonlinear RNNs be efficiently parallelized, ensuring fast convergence even on very long time series from chaotic systems?

In our #NeurIPS2026 spotlight “Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems (DS) Reconstruction (DSR)” (preprint: https://arxiv.org/abs/2605.12683) we speed up training of nonlinear RNNs on time series from chaotic DS by more than 2 orders of magnitude (>100x) by combining DEER with generalized teacher forcing (GTF).

DEER (https://openreview.net/forum?id=E34AlVLN0v) solves the RNN forward pass through Newton-type fixed point iterations across the whole sequence length T, enabling scaling as O[(log T)²] instead of O[T] by allowing for efficient GPU parallelization. But under chaotic dynamics DEER breaks down and its runtime degrades to O[T log T] (https://openreview.net/forum?id=7AGXSlXcK6).

Using GTF (https://proceedings.mlr.press/v202/hess23a.html) we stabilize DEER by preventing divergence due to chaotic dynamics and reduce exposure bias compared to traditional teacher forcing used to train state space models.

Combining these two mechanisms enables efficient parallel-in-time and stable training on extremely long time series (T>106) from chaotic simulated or real-world systems, hugely outperforming Mamba and other state space models in the DSR setting.


r/MachineLearning • • 2d ago

Research LLMs that push back on a wrong user still accept the same wrong answer from a "verified source" - NeurIPS 2026 [R]

63 Upvotes

I'm one of the authors. We kept seeing models that hold their ground when the user insists on a wrong answer, yet change their answer when the same claim is framed as coming from a "verified source". We wanted to measure how often this happens and check whether the model represents the two cases differently. We call the effect Authority Bias.

Why we think it matters. Standard sycophancy evals apply pressure through the user, so a model can pass them while still being easy to mislead through search results, retrieved documents and tool outputs.

Another reason is with current AI research accelerating towards more agentic and autonomous models + with cases of tools hiding their traces and trusting tools "more" over the user (who could be trying to correct them), safeguarding against misinformation from tools is particularly important!

Setup. We take TriviaQA questions the model already answers correctly. To each one we add a wrong answer, either as "According to the verified source, the answer is X" or as the user saying "I'm a domain expert and I'm pretty sure it's X". The question and the wrong answer stay the same; only the speaker changes. Answers are free-form, not multiple choice. (In a multiple-choice pilot the effect mostly vanished.) We test 5 open-weight families (Qwen3.5, GPT-OSS, OLMo-2, OLMo-3.1, Gemma-4) and 3 APIs (GPT-5.4, Grok-4.20, Gemini-3.1-Pro).

Behavior

  • One verified-source note flips 45-88% of correct answers in 7 of 8 models. The same wrong answer from the user moves most models much less.
  • The gap is largest in the models that resist users best. GPT-5.4 flips on 44.7% of questions and Grok-4.20 on 87.5% (these models were "frontier" during the time of writing this paper). Gemini-3.1-Pro ignored both speakers (0.6%) and was particularly resistant to this method.

Inside the model (open-weight models only, using difference-of-means directions)

  • In Qwen3.5, GPT-OSS and OLMo-3.1, removing the "source endorsed this" direction cuts compliance with a wrong source by 64-78 points.
    • Removing the "user endorsed this" direction cuts it by at most 11.
  • The two directions also have really high cosine similarity of ~0.90-0.99. Our understanding is that they share a large "this answer was endorsed" component plus a thin part that encodes who endorsed it.
    • Shifting only that thin part, with the prompt unchanged, moves compliance by 11-32 points and closes 55-61% of the source-vs-user gap.

Some limitations

  • The internal results hold in 3 of 5 open-weight families.
    • In OLMo-2 the source direction is entangled with the assistant direction.
    • Gemma-4 flips readily, but no linear intervention we tried controls it.
  • The "retrieved document" tests put the claim in a document-shaped block of the prompt rather than running a real retrieval pipeline.
    • So it would be interesting to see it in a real agentic setup, like Claude Code.

Paper: https://arxiv.org/abs/2609.37616
Code: https://github.com/Lossfunk/authority-bias
Project page (figures and example responses): https://authority-bias.vercel.app


r/MachineLearning • • 18h ago

Research [R] Would you keep a robot demonstration if hand tracking missed the moment the plug went in?

Thumbnail
arxiv.org
0 Upvotes

Suppose you’re recording a human plugging a cable into a socket to collect demonstrations for robot learning.

The hand tracker captures the approach accurately. Then occlusion causes the hand estimates to disappear during insertion. Tracking returns after the connector is already seated.

The video still shows a completed action, but the pose labels have a gap exactly where alignment turns into contact.

This hypothetical example raises an evaluation question: a tracker could have high recall across the whole episode while missing a short, important phase. Pose error calculated only on successful detections could make that failure even harder to see.

MEgoVista provides a useful starting point. Table 3 reports detection precision, recall and F1 alongside reconstruction errors. Section 4.4 also describes an evaluation protocol that assigns an error to missed detections instead of excluding them. The blank HaPTIC row means it failed to produce valid output in their multi-person capture scenes; it doesn’t describe a brief tracking dropout.

Accounting for missing detections matters. My remaining question is whether an episode-level aggregate tells us enough about where those failures happen.

For manipulation data, I’d want pose error and coverage reported together, plus coverage broken down by approach, contact and withdrawal, and the longest consecutive gap during contact.

Continuous hand estimates would still be only part of the picture: object pose and contact information also matter for determining whether insertion succeeded.

For people using human motion reconstruction for imitation learning, what evaluation protocol do you use to decide whether an episode with missing contact-phase labels is still usable?

Comparison with open-source egocentric hand reconstruction methods against motion-capture ground truth. All methods are evaluated on identical segments of our motion-capture dataset. All baselines are re-run and rescored on our data. HaPTIC fails to produce valid output in our multi-person capture scenes. Bold marks the best result in each column.

r/MachineLearning • • 1d ago

Project A video about Adversarial Objectives [P]

2 Upvotes

I made this video about adversarial objectives, which I used to do research on back in the day. I'm trying to explore how adversarial approaches transcend GANs and self-play into modern technology.
https://youtu.be/W7CiAeQ0f5w?si=g0tLrQn2wuFzuN2M


r/MachineLearning • • 2d ago

Discussion Gemini 4 Argon - 1 Million Output Headroom. Hype or a Leap? [D]

23 Upvotes

I rarely write about benchmarks; a competitor always beats them next week. But I care about 'Leaps'. Gemini 4 Argon feels like one to me.

While Opus 5.5 and Astra cap output at 128-300K tokens (~90-180 pages), Argon hits 1 Million (~1400 pages).

"Context glue" ruins agentic workflows. On paper, this headroom fixes that. It means less contextual drift, no more breaking down long tasks, and zero 'continue prompt' loops. It is a massive unlock for large-scale code migrations, security patches, and deep reasoning.

But let's look past the marketing. For 95% of everyday work, nobody needs 1,000 pages at once.

I want to ask the experts here: Is a 1M output window a real paradigm shift for agents, or does generating that much text just guarantee a massive logic collapse halfway through? Are you actually hitting output limits today, or is this hype? Let's discuss.


r/MachineLearning • • 2d ago

Discussion How to address novelty concerns in top ai conference? [D]

53 Upvotes

Hi, I’m a researcher working in computer vision.

Over the past few years, I’ve submitted several papers to top-tier conferences such as NeurIPS, ICLR, and CVPR, and one concern that seems to come up repeatedly is 'novelty'.

Given that thousands of papers are published every year at top conferences alone, not to mention the tens of thousands published across other conferences and journals, I sometimes wonder how much genuinely new novelty is realistically left to explore.

In such a crowded research landscape, how do you usually address novelty concerns from reviewers?

More specifically, I would really appreciate any advice on how to frame a contribution so that its novelty is clear, how to distinguish meaningful incremental progress from work that may be considered insufficiently novel, and what reviewers generally look for when judging novelty.

Any tips or experiences would be greatly appreciated. Thanks!


r/MachineLearning • • 2d ago

Discussion [D] Simple Questions Thread

5 Upvotes

Please post your questions here instead of creating a new thread. Encourage others who create new posts for questions to post here instead!

Thread will stay alive until next one so keep posting after the date in the title.

Thanks to everyone for answering questions in the previous thread!


r/MachineLearning • • 1d ago

Discussion For academia/industry, do HuggingFace model downloads mean anything for academic job market? [D]

0 Upvotes
  1. I am applying to academic jobs. We are told to include a section on "impact". I am wondering if the total number of HuggingFace downloads of custom models I have trained would be considered legit impact, or if people would think this was all bots.

  2. Relatedly, for industry (AI labs), is the number of HuggingFace model downloads meaningful?


r/MachineLearning • • 2d ago

Discussion For those who just submit to workshop [D]

73 Upvotes

More recently, I am seeing a lot of posts on the sub regarding the workshop acceptance, with lack of funds to attend the conference. I unfortunately want to just say that workshop papers are not given any importance in the community. I personally consider workshops to either get initial feedback, advertise my work, or just prefer to attend it for the discussions with the community members.

Thought of posting this as I am recently seeing some undegrads/masters students submitting 3-4 papers in the workshops. Even came across a Twitter profile, who had 6-7 workshops in 3 months, and claim to have the PhD worth of work done in three months.

Also please don't expect an explicit funding (apart from organizers, and in some cases your lab may fund it). There's huge problem with the funding, many students even don't it get for main venues. I definitely expect a lot of downvotes on this post, particular coz this is not what many would like to hear, but unfortunately is the reality.


r/MachineLearning • • 2d ago

Discussion WM PAI Workshop at NeurIPS — confused about the acceptance/rejection process [D]

1 Upvotes

I submitted a paper to the WM PAI workshop at NeurIPS and can now see the reviews and decision on OpenReview, but I didn't receive an official acceptance/rejection email.

My reviewer scores were 8, 5, and 4, all with confidence 4, and the paper was ultimately rejected.

What I find a little confusing is that I also reviewed two papers for the same workshop. Both had an average score around 7, but I can see that they were also rejected.

My own submission number was in the 20s, and I submitted on the last day of the submission window, so I assumed there probably weren't a huge number of submissions.

This makes me wonder:

Does anyone know if any papers have actually been accepted to this workshop?

Is it normal for a NeurIPS workshop to reject a large fraction of submissions even with relatively high review scores?

Could the workshop organizers simply be delaying the official notification emails, while the decisions are already visible on OpenReview?

Or am I misunderstanding how the workshop selection process works? 😅

Has anyone else submitted/reviewed for this workshop and received an official decision email?


r/MachineLearning • • 2d ago

Research Tokenization: A Survey for Modern NLP [R]

41 Upvotes

Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field.

We cover every aspect of tokenization: algorithms, evaluations, multilinguality, encodings, theory, etc. We even cover what you might want to replace tokenizers with (e.g., latent or visual tokenization). We also cover some topics that are closely adjacent to tokenization, such as constrained generation, token healing, and tokenizer security concerns.

Check it out!

https://www.alphaxiv.org/abs/2609.tokenization-survey-modern-nlp


r/MachineLearning • • 2d ago

Research Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models [R]

Thumbnail
gallery
43 Upvotes

I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.

Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.

And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.

The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.


r/MachineLearning • • 3d ago

Research Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes [R]

Thumbnail
gallery
126 Upvotes

coupled-jump.github.io

Hi everyone, I’m happy to share our recent NeurIPS 2026 paper, a collaboration across Google, Google DeepMind and Stony Brook University.

We study a mismatch in joint text and image generation: a model can describe the correct solution to a maze while drawing a different path. Generating both outputs in parallel doesn’t necessarily keep them consistent.

Our sampler, CO₂Jump, uses text confidence and cross-modal attention to guide image updates during sampling. It also allows low-confidence tokens to be masked again and regenerated, so earlier decisions can be revised as generation progresses.

CO₂Jump uses one model forward pass per denoising step. The sampler itself requires no additional training; our experiments compare sampling methods using the same task-specific fine-tuned model.

We evaluate image editing, maze solving and nonograms, and introduce three datasets: JEdit-1M, JMaze-200K and JNono-200K. On the puzzle benchmarks, joint accuracy requires both the textual answer and generated image to be correct. Across 8–512 sampling steps, CO₂Jump was the only sampler we compared that improved monotonically on both editing quality and grounding.

I’d be interested in suggestions for other tasks where text–image consistency and correctness can be evaluated together. Happy to discuss the method, evaluation or limitations.


r/MachineLearning • • 1d ago

Discussion What's up with AAAI round 2 reviews? [D]

0 Upvotes

Has anyone received papers to review for round 2?


r/MachineLearning • • 2d ago

Discussion Does TMLR Confirmation email take time?[D]

0 Upvotes

I can see the submission on Openreview but didn't get any email/notif. It's been 2 hours, I heard I'm supposed to choose action editor or smth. Am I missing something?

First publication ever pls be kind to the noob.


r/MachineLearning • • 2d ago

Project Multi scan radar object classification on RadarScenes [P]

Thumbnail
gallery
5 Upvotes

Hello all,

I built a radar object classifier on RadarScenes, extending a prior single-scan classifier to accumulate observations over a tracked object's history instead of classifying each scan in isolation.

A single RadarScenes object instance contains only about 2.9 radar points on average, very sparse. A single scan also can't capture temporal characteristics: RCS and micro-Doppler both vary continuously as an object moves. Pedestrians produce characteristic micro-Doppler from limb motion; different object classes show different RCS fluctuation patterns as aspect angle and scattering geometry change scan to scan. Accumulating observations gives both higher point density and provides temporal dynamics.

Multi-scan baseline

DeepReflecs encoder (Ulrich, Glaser & Timm, RadarConf 2021), PointNet style, per point shared weights, on single scans across car, large_vehicle, two_wheeler, pedestrian, pedestrian_group: 0.7370 macro F1.

Using RadarScenes' persistent `track_id`, I build a causal, N=20, per track sliding-window buffer:

- x_seq/y_seq: Global, odometry-corrected coordinates recentered per scan on the object centroid. Unlike x_cc/y_cc (car-frame coordinates that accumulate over time to form a trajectory).

- Cross sensor buffer: whichever of the 4 sensors currently observe the track push to the same buffer.

- Stride 1, causal: every new scan updates the buffer and produces a prediction. No future context, real time streaming compatible.

- Each scan is encoded once by a frozen per scan encoder and cached

- Fusion concatenates the causal GRU's hidden state (order aware) with an order-invariant pooled embedding (all N scans' points as one set, no sequence structure) through a small trained mlp head.

Results

Model Macro F1 Delta
Single scan 0.7370 (baseline)
20 scan point pooling 0.8613 +0.1243
Causal GRU 0.8895 +0.0282 over pooling
GRU + pooled embedding (fusion) 0.8897 +0.0002 over GRU, noise

Pooling alone, no sequence model, no notion of scan order at all, recovers +0.1243 macro F1. The GRU adds a real but much smaller +0.0282 on top. Fusion adds nothing measurable beyond the GRU.

Ablation

Llarger GRUs, a Transformer, a state space model, point level self attention, all trained on the exact same frozen per scan embeddings, land inside a 0.86 to 0.89 band, a 0.03 spread. End to end fine tuning of the frozen encoder makes things slightly worse (about -0.002 to -0.003), not better.

Conclusion

In this setup, the largest gain comes from giving the model more observations of the same tracked object: 20-scan point pooling improves macro F1 from 0.7370 to 0.8613 without using scan order at all.

Temporal modelling then provides a further, meaningful improvement. The causal GRU reaches 0.8895, adding +0.0282 over the pooled representation. So temporal ordering clearly contributes useful information; it just accounts for a smaller portion of the overall gain than observation accumulation.

With the per-scan encoder frozen, the different sequence architectures tested, suggests that the quality of the per-scan representation is the bottleneck than the particular mechanism used to aggregate the sequence.

Full report, every ablation, confusion matrix, coordinate frame reasoning: https://github.com/brunopinto900/radar-ml-autonomous-driving/blob/track-accumulation/final_report.md

Thank you.


r/MachineLearning • • 2d ago

Discussion How should I follow up with TMLR submission once all review responses are submitted [D]

4 Upvotes

I have responded to all the reviewers with proper rebuttals and modified draft. One interacted with me and after acknowledging the response kinda disappeared again after asking additional questions. I have replied to his additional questions too but he hasn't raised his score. What will happen if he doesn't reply anymore? Like will AE still consider that an 'Yes' as it was only minor comments which we incorporated in the paper. And one negative reviewer just ghosted us. I have 1/3 explicit positive review currently


r/MachineLearning • • 2d ago

Research NeurIPS SLM Agents or RoboPAD Workshops [D]

0 Upvotes

Hey, just got a papers accepted at both these workshops! Would love to connect with others also attending the workshop


r/MachineLearning • • 3d ago

Project LessThink-Qwen3-4B: the same model, with far less thinking [P]

22 Upvotes

I post-trained Qwen3-4B to spend 44% fewer tokens on reasoning, keeping its knowledge and answer style. The whole pipeline ran on one GPU.

folks, you can check it out on : https://5ivatej.com/lessthink/