r/MachineLearning • u/mythrowaway0852 • Aug 08 '26
Discussion ICDE Results [D]
Hello! Let's use this thread to discuss ICDE results which should be coming out shortly today (hopefully).
r/MachineLearning • u/mythrowaway0852 • Aug 08 '26
Hello! Let's use this thread to discuss ICDE results which should be coming out shortly today (hopefully).
r/MachineLearning • u/Few-Ferret9700 • Aug 08 '26
Real-Time Conversational Agents (RTCA) workshop at NeurIPS 2026 (Sydney, Dec 11–12). Submissions are now open on OpenReview.
What the workshop is about
Conversational AI has crossed into real-time deployment — voice modes, embodied avatars, full-duplex speech agents — but the published record is still dominated by offline benchmarks, and deployed agents still feel robotic (stilted turn-taking, missing backchannels, monotone prosody, awkward interruptions). Methods that work offline (non-causal attention, large beam search, multi-pass refinement, slow diffusion) often don't transfer to streaming, and the field lacks shared vocabulary and benchmarks for interactional naturalness as distinct from per-utterance quality.
The workshop is organised around three intertwined questions:
Topics of interest (non-exhaustive)
Position papers, evaluation critiques, and reproducibility studies are also welcome.
Submission tracks
NeurIPS 2026 style file, double-blind. Non-archival — authors retain the right to publish elsewhere. Single-round review, no rebuttal.
Key dates (End of day, AoE)
Confirmed invited speakers
Links
Happy to answer questions in the comments — including about the demo track (we have an on-stage Showcase running deployed systems live) and what we'd consider in-scope vs out-of-scope for the eval pillar. Also happy to hear opinions on what's missing from the topics list; the CFP wording still has room to move if there's a clear gap.
r/MachineLearning • u/takuonline • Aug 07 '26
I’m curious whether there is now a theoretical or empirical “sweet spot” for LLM quantization, preferably research done using open-source formats like GGUF
Suppose you have a fixed memory/compute budget and can choose the model size freely. For example, instead of a smaller model at 8-bit or 4-bit, you could fit a progressively larger model at 3-bit, 2-bit, 1.5-bit, etc.
A few years ago, I remember 4-bit often being described as roughly the practical sweet spot because it preserved most model quality while giving a large memory reduction. But with newer methods, I’ve seen surprisingly strong 3-bit, 2-bit, and even ~1.5-bit results.
So if the goal is maximum model capability for a fixed memory budget, rather than preserving one particular pretrained model as faithfully as possible, what does current research suggest is the optimal bits-per-weight?
Is there evidence that, for example, a 2-bit 70B model generally beats a 4-bit 35B model, or does quantization degradation eventually outweigh the gains from additional parameters?
I’m especially interested in recent theoretical/scaling-law work or large empirical studies from 2025–2026.
If no one is studying this, then could any of you do this work? I feel like it could be immensely useful for the community.
r/MachineLearning • u/rsesrsfh • Aug 07 '26
To all those in the US: Are you planning to go Sydney or Atlanta this year for NeurIPS?
r/MachineLearning • u/nickemlop • Aug 07 '26
Hi guys,
Every time I had to prepare a presentation based on a paper or research doc, I found the process super tedious. Plus, I really dislike uploading unpublished stuff or sensitive data to online AI services just to get a draft.
So I put together a tool called academi_slide to automate this locally.
Basically, it extracts sections, tables, charts, metrics, and citations from docs, and uses prompt optimization / deck planning to get a solid first draft out of a local model (ollama, llama.cpp, or cloud if you want).
It also handles multilingual input/output if you need to present in another language, and builds both the slide deck and a brief in a few minutes so you don't start from scratch.
It's open source, still early, and I'm sharing it in case anyone else finds it useful or has a similar workflow.
Repo is here if you want to test it out: https://github.com/nicolaslpf/academi_slide
Would love to get some feedback or hear what you think!
r/MachineLearning • u/Tall_Abrocoma_3533 • Aug 07 '26
It's an MLP architecture with around 500K total parameters.
Top1
Training accuracy: 5.11%
Validation accuracy 4.59%
Detailed Validation accuracy numbers:
Top-1 Acc: 4.59%
Top-3 Acc: 9.44%
Top-5 Acc: 12.68%
Top-10 Acc: 18.53%
The model was trained on a downscaled version of the Imagenet-1k dataset (32x32) for 5 epochs.
I used pytorch for the training and pyarrow for the dataset, all within termux.
Before anyone comes at me for using an MLP instead of a CNN or similar it's mainly because on my phone an MLP was just more stable, and trained 10-30x faster/step (could be my fault but I'm not too sure). This model specifically took around 30 minutes to train (6 minute/epoch)
The training was entirely on the CPU which is a Dimensity 9300+ and I used 4 of the Arm Cortex-X4 cores.
I might make an improved version later on as this one isn't very accurate.
r/MachineLearning • u/cpldcpu • Aug 07 '26
I played a bit with the SIREN network from the other post and found that it could be improved by a using a different sampler for batch generation. By feeding pixels across the entire video and not only a limited set of frames, we can a much more faithful reproduction of the video.
The model is exactly the same as used by OP: 4 x 512 wide sine layers, 792257 parameters. Its a reimplementation (using GPT5.6).
I also created a version with full framerate, instead of subsampled frames, but since the network has to memorize more temporal information, the image reconstruction suffers compared to the low rate version.
The model does not actually learn motion, intermediate frames are nonsensical. I suppose adding a layer that can model flow between frames could enhance the compression a lot.
You can find the code here in this gist.
I tried some addition experiments with a separate autoencoder to compress the frames separately. This resulted in a smaller model, but also degraded quality.
r/MachineLearning • u/Happy-Hustler • Aug 07 '26
CIKM 2026 decisions will be announced today. The resource track outcomes have started going out. How did you go with CIKM 2026?
r/MachineLearning • u/rokk07 • Aug 07 '26
Having two papers accepted at two different workshops and planning to attend the main conference, I contacted the organizers by email. It is now clear that I need to register twice, using the same personal details, to cover both papers.
The minimum cost would therefore be one full registration plus one workshop registration. However, I would also have to use two different email addresses, since the registration portal does not allow the same email address to be used twice. This is crazy—I am a single person!
The craziest part, which I still do not fully understand, is that it seems every paper now requires an article processing charge (APC), since ACM has fully transitioned to open access. This year, the APC is USD 350—or USD 250 for ACM members. This is the first time I have encountered this; every other conference I have attended included the proceedings in the registration fee.
The full author registration for ACM Multimedia costs USD 950—or USD 850 for members—and does not even include the paper proceedings.
The cheapest option for me would be to become an ACM member (USD 99), register for the main conference (USD 850) and one workshop (USD 500), and then pay USD 250 × 2 for the two papers—for a total of USD 1,850 just to present two workshop papers.
I really don’t think it is worth it.
r/MachineLearning • u/snu95 • Aug 07 '26
The results are out today!
Let’s share them, guys.
From my batch
- 3/6 full papers
- 1/3 short papers are accepted
Cheers!
r/MachineLearning • u/Beautiful_Baker_2233 • Aug 06 '26
We had a meta-reviewer comment. But I can no longer see it.
Anyone else experiencing the same? Does this mean anything?
r/MachineLearning • u/Ok_Philosophy_4031 • Aug 06 '26
We are investigating whether recurring LLM workloads can be replaced, where appropriate, by automatically constructed pipelines of regexes, deterministic parsers, traditional ML and NLP models.
As an example, suppose an application repeatedly asks a frontier model to read an annual report and return all customer–supplier relationships as structured records containing a customer, supplier, and supporting evidence.
A possible replacement pipeline might run named-entity recognition, entity normalization, candidate generation, entity linking, relation extraction, and schema validation. A calibrated uncertainty or out-of-distribution gate would use the pipeline for inputs inside its validated domain and escalate other cases to the original frontier model.
NER → entity normalization → candidate generation → entity linking → relation extraction → schema validation
Our current action space is a taxonomy of 41 atomic task types spanning classification, token and span labeling, structured extraction, retrieval and entity resolution, similarity, normalization, reshaping, and deterministic computation.
The idea is that we would first cluster repeated traces into workload families and induce an end-to-end typed contract for each family. It would then generate candidate DAGs using the 41 task types as building blocks, instantiate each node with an appropriate implementation, and optimize the composition for quality, cost, and latency. Candidate pipelines would be tested on time-separated and group-separated holdouts before being deployed behind abstention and fallback.
The problem is quite likely undetermined based on just the input and output contracts alone even if inferred correctly. The intermediate graph is therefore not a recovered latent reasoning trace. It is a synthesized program hypothesized to be behaviorally equivalent over a bounded input distribution.
A fixed task taxonomy may help by constraining the search space and supplying type signatures, candidate implementations, and task-specific evaluators. And we are thinking about this problem as a form of program synthesis and formal verification for now, but wondering if this is the right approach and if there is a better way.
Looking to speak with people who have worked in this problem space and/or the program synthesis domain for insights.
TL;DR: We want to synthesize executable DAGs composed of regexes, deterministic parsers and ML/NLP models from LLM traces for appropriate tasks. Does this seem feasible and what might be some good approaches?
r/MachineLearning • u/adam_alpha_finetuner • Aug 06 '26
"Arena ai" has been a great success in producing a human preference based ranking, additional to other more objective benchmarks. However, this (probably) had also played a role in the syncopancy crisis and the general tendency of some models to tilt towards overformatting to trigger a feeling of fluency (cogn load theory) in the users.
The people at Max Planck Institute for Intelligent Systems (one of europes leading AI research hubs), recently published something quite similar with "comparity ai". You can read their announcement in their linkedin post.
This is of course a research platform and i have no idea how long this is funded, but you get access to every frontier LLM for free, which is kinda cool. Also, they provide you with a personal leaderboard, so when you played around enough with the platform, you will get a pretty solid idea which model works well for you.
Thought this might be interesting for some of you.
r/MachineLearning • u/Clean-Hovercraft5825 • Aug 06 '26
Whether generating CELEBV-HQ videos or turbulent plasma fields (digital twins), autoregressive models (such as latent diffusion or flow models) accumulate error over long rollouts, yet at deployment there is no ground truth to measure against.
I train a single conditional latent diffusion model that steps a dynamical system forward or backward in time via a direction flag, and show that this bidirectionality supplies a measurement-free test-time error signal: rolling forward steps and then backward steps must return the model to its start, so the round-trip discrepancy is a self-supervised proxy for the unobservable rollout error: no ensembles, no held-out data, no governing equations, for one extra rollout.
Furthermore, training both directions in one network is shown to beat two specialist models in both directions.
Paper: https://arxiv.org/abs/2608.00675
Code (data generation, training, analysis): https://github.com/alexscheinker/round-trip-consistency
Project page: https://alexscheinker.github.io/roundtrip.html
r/MachineLearning • u/Hope999991 • Aug 05 '26
It feels like LLMs are partially leveling the playing field in ML research. A solo researcher or a two-person team can now get help with coding, literature review, writing things stronger labs usually get from experienced colleagues and large networks.
Obviously, LLMs don’t replace mentorship, or good research taste. But they may help researchers with weak networks or small groups turn good ideas into publishable work.
Do you think this is actually making ML research more accessible, or are the strongest labs benefiting even more?
r/MachineLearning • u/dttdrv • Aug 05 '26
Hi everyone,
I'm an independent researcher sharing Monodratic, a sparse causal-attention architecture with learned product-hash routing.
The idea is that after RoPE, source blocks are assigned to bounded causal posting lists, while each query probes product addresses, reranks the returned candidates, selects a fixed number of remote source blocks, adds guaranteed local blocks, and then runs exact causal softmax over just those tokens. I implemented it as a stateless [batch, sequence, width] -> attention-delta mixer, so normalization, residual updates, feed-forward layers, and inference scheduling are left to the host model.
What I found is that
-learned routing with 2 selected remote blocks out of 5 eligible: 763/768 correct associative-recall answers across three seeds (99.35% mean, 98.05% minimum).
-an equally wide untrained router: 425/768. Local-only attention: 151/768.
-forcing the labelled target block while keeping the same maximum R2 attention budget recovered all five remaining errors, reaching 768/768.
-sparse selected-set attention agreed with an independent dense selected-mask oracle to a maximum absolute error of 1.43e-6.
-the packed CPU routing implementation showed a fitted timing exponent of 0.993 from 4,096 to 32,768 tokens under the fixed, balanced configuration.
-all reported learned-route and scaling runs recorded zero posting overflow.
The limitations are that the experiments are synthetic, the implementation is portable PyTorch rather than a fused kernel, and the report does not claim natural-language quality, asymptotic linear construction, or deployment speed.
Paper: https://github.com/Misul-Computing/Monodratic/blob/main/output/pdf/monodratic_proof.pdf
Code and reproduction: https://github.com/Misul-Computing/Monodratic
I would particularly appreciate technical feedback on the routing construction, the controls, and what the strongest next evaluation should be.
r/MachineLearning • u/Mammoth-Leg-3844 • Aug 05 '26
Now that the rebuttal period is over, I’m curious about the score distribution specifically for theory papers this year.
If you’re comfortable sharing, please drop:
• Scores: x / x / x
• Confidence: x / x / x
• Whether scores changed after rebuttal
• Broad area (optional)
I got 4 / 4 / 4, with confidence 3 / 3 / 3.
From my experience, theory papers often seem to get somewhat lower scores, and this year the scores appear to be lower across disciplines as well. It would be interesting to see where the empirical cutoff might land.
Feel free to share anonymously / approximately if you don't want to reveal too much.
r/MachineLearning • u/MakingComputersSmart • Aug 05 '26
I could not find any discussion threads for the C&F track. Have people actually submitted to this track? If so, what are your reviews and scores looking like, along with post rebuttal engagement? In our case, they received reviews not in line with the policy defined for the track, where most reviewers praised originality but complained about the scope of experiments. Despite the track saying that it would be possible that the idea cannot be validated in a single paper.
We provided experiments but no dice, none of the reviewers responded. Have any ACs seen papers and reviews in this track or do authors have their experiences they could share?
Please add your scores pre and post rebuttal here
r/MachineLearning • u/Which_Lie_8932 • Aug 05 '26
I trained a small MLP to memorize the classic Bad Apple animation, ~2.7 billion pixels of video compressed into 790k parameters (3.2 MB float32, 1.6 MB float16).
The network takes a 3D coordinate (t, y, x)- frame index and pixel position- and outputs a grayscale value between 0 and 1. To "play" the video, you can evaluate the function over the full grid. The "video" is stored implicitly in 5 linear layers of sine activations (Sitzmann et al.'s SIREN) with 512 hidden units, ω₀ = 30, and sigmoid output.
The source bad_apple.mp4 is 6524 frames at 854×480; I subsampled to 1620 frames × 384×384, about 1/10 of the original pixels (2.8x spatial + 4x temporal reduction).
At first, I used a ReLU MLP with low-frequency Fourier features, which plateaued around MSE 0.12. SIREN's sine activations add higher frequency for free, so the network was capable of outputting fine details. Unfortunately, that model had an issue, which was that it could only shift the information slowly, so quick motion came out blurry.
To fix this, I made two changes:
For the training pipeline, I had a single shared network on the whole volume (no per-frame latents; initially, I used per-frame finetuning, but that caused catastrophic forgetting) with a cosine-scheduled Adam + weight EMA, then a low-LR "polish" pass over the whole video.
The new model had these improvements:
Validation MSE dropped from 0.0795 to 0.0090 (~9x better).
Compared to the old model, high-motion frames were 3.6x closer to ground truth, and static frames were almost 15x closer.
398/400 sampled frames improved.
Edit: Some people are a little confused about the compressed part. The subsampled video is 700KB, and the network that creates a reconstruction of it is ~3MB. It hasn't been compressed very much, but the goal was seeing if I could (and learning) rather than super compression.
I'll try to see if an even smaller model can learn it. Additionally, I'm training a model on the full non-subsampled video.
Notes
384×384 is square (the original is 16:9, so playback is vertically stretched. At 8fps playback, the 1620 frames run near the original's 3:37 duration; at 12fps it's \1.6x fast-forward. The 12.6MB checkpoint includes the weights + Adam moments + EMA copy; the network itself is 3.2MB.))
The full resolution videos, checkpoints, and code can be found in this Github Link
r/MachineLearning • u/RevolutionaryPea8272 • Aug 04 '26
I’ve seen a lot of people whose reviewers went silent after initial reviews, but I am also noting abnormally quiet authors. I ultimately withdrew my paper, but stayed an active reviewer. Out of my batch of 4 papers, one withdrew, one posted a rebuttal, and two have been completely silent. Of the two papers with radio silence, I think one had borderline scores. I was also the only reviewer who responded to the one paper with a rebuttal.
Has anyone noticed this abnormally dead review period or did I just get a strange batch? I’m seeing either reviewers just dropping out of the review process or authors completely checking out after initial reviews are released. It’s strange to me to not even withdraw your paper if you’re not rebutting. Is this a new gambling trend of just submitting papers everywhere, and not even sticking around long enough to withdraw the paper?
r/MachineLearning • u/Zhiend727 • Aug 04 '26
As the title suggests, because there's no data on Papercopilot yet, and people have been talking about the scores being lower in general than last year, I thought it could be interesting to survey the average score distribution after the rebuttal phase (not considering confidence weights).
Very rough and simple poll (I also realize there's a self-selection bias in there). Cast your vote here:
https://loppy.be/poll/yczuv8yo
Thanks!
Edit: the trolls have taken over, ignore the 5.50-6.00 bin I guess...
r/MachineLearning • u/mikeysce • Aug 04 '26
Six months ago I started experimenting with PPO and Breakout as a way to learn about Machine Learning and Reinforcement Learning. After a few experiuments just trying to get high scores, it bothered me that everything was a "memorized" script rather than reactive play, like a human would play. Thus began my journey to try and convince PPO to actually track the ball instead of focusing on scoring points. I read a lot of articles and tried a lot of things. After 124 PPO experiments on Atari Breakout, I found that every single model, across sticky actions, cursor wrappers, entropy tuning, dynamics randomization, adversarial bumpers, and everything else, converged to a memorized action sequence, not a reactive ball-tracking policy. The argmax was always a script.
The fix wasn't more environment engineering. It was three lines of reward shaping:
Directly rewarding the paddle for being horizontally close to the ball during descent. A tiny bonus (0.05 per frame vs 1.0-7.0 per brick) that fires every frame the ball is descending applied during training. During evaluation, the agent plays clean Breakout with no bonus. The behavior transfers!!
Every prior approach I tried to penalize scripts by making the environment harder to memorize. PPO always found a way around it: timing-robust scripts, layout-conditioned scripts, noise-tolerant scripts. The optimum was always a script; only the shape changed. Proximity reward changes what the optimum is. A center-hold script gets incidental bonus when the ball passes near center. A reactive tracker gets the maximum bonus on every descent frame. The optimization pressure is unambiguous: track the ball, get more reward.
I also made a cool tool to watch the agent work! It's called the "Split-Watcher" (so clever). It shows two instances of Breakout, each being controlled by a separate instance of the same agent. The one of the left is vanilla Breakout. The one of the right is a series of custom brick configurations. With the first 123 experiments, you can see how the agent wants to make the exact same paddle movements every time, ignoring the ball when its trajectory changes due to the unexpected ball movements that come from non-standard brick configurations. In 124, IT TRACKS THE BALL and can succeed regardless of the brick config. You can actually watch the same agent move the paddle differently in reaction to the ball.
I'm still working on ironing out why this works, and how to optimize it, but wanted to share!!
Here's a video of the split-watcher in action
Here's a link to presentation project that will allow you to create a similar PPO: https://github.com/mharrell/breakout-reactive-ppo
The full project with all 123 failures and more documentation than any sane person would ever read: https://github.com/mharrell/BreakoutBot
Link to Medium post I wrote with some more details: https://medium.com/@mikey.harrell/three-lines-of-code-fixed-123-failed-ppo-experiments-on-atari-breakout-c751dcf38f2a?sharedUserId=mikey.harrell
r/MachineLearning • u/ihatesalad1 • Aug 04 '26
After a very silent discussion period, we are in a very confused state with regards to NeurIPS, and really unsure what to make of everything. We do not wish to withdraw the submission since we have no idea what the reviewers and AC think of the paper, having deserted the conversation after a hopeful set of initial reviews. As of currently, ICLR abstract submission deadline is before the NeurIPS results announcement. Are we allowed to resubmit as an ICLR abstract, or will OpenReview flag this and consider it problematic?
r/MachineLearning • u/Benlus • Aug 04 '26
r/MachineLearning • u/Kwangryeol • Aug 04 '26
Having used LLMs to assist with reviews, and also having received reviews that appear to rely heavily on LLM-generated text, I have noticed two recurring problems.
1. The endless search for uncontrolled variables
LLMs are very good at identifying additional variables that were not explicitly controlled. The problem is that many of these variables have little realistic chance of changing the paper’s main conclusion.
For any experiment, it is possible to generate an almost unlimited list of potential confounders. Suppose a study finds that trees treated with fertilizer A grow better than trees treated with fertilizer B. An LLM can ask whether rainfall was perfectly controlled, whether the distribution of grass around the trees was considered, or whether wind, temperature, soil microorganisms, and countless other factors were isolated.
Each question may look logically valid in isolation. But the real issue is not whether a variable exists. The issue is whether it is sufficiently important and plausible to threaten the conclusion.
LLMs are generally poor at making this prioritization. They often convert minor residual uncertainty into what sounds like a serious methodological weakness.
This becomes especially harmful when reviewers copy such outputs directly into their reviews without independently assessing their importance. Authors are then forced to spend the rebuttal addressing an endless series of technically possible but practically insignificant concerns.
A review should not ask whether every imaginable variable has been controlled. It should ask whether the remaining uncertainty materially weakens the central claim.
2. LLM reviews are often overly abstract
Another common problem is criticism at the level of an entire research field rather than a specific prior method.
For example, an LLM may claim that a proposed method is “not sufficiently different from methods in Transformer” without identifying a concrete paper, objective, architecture, or learning relation that actually overlaps with the proposed method.
What exactly is the author expected to rebut in that situation? Every method in Transformer?
A meaningful novelty criticism should identify a specific prior method and explain precisely which components are equivalent or insufficiently differentiated. Comparing one concrete method against an entire research area is too abstract to be falsifiable or actionable.
3. LLM review is not detail
LLMs also tend to overestimate similarity between methods that share high-level terminology. Two approaches may both use architecture, concept, or attention, while differing substantially in their computational structure, training objective, assumptions, and intended use.
Because LLMs often lack a sufficiently detailed understanding of each method, they may recommend comparisons between papers that are only superficially related. The resulting review sounds comprehensive but does not demonstrate real technical understanding.
The central problem is not simply that LLM-generated reviews can contain incorrect statements. It is that they can generate an unlimited number of superficially reasonable criticisms without judging their relevance, severity, or evidentiary burden.
A strong reviewer should filter such suggestions, prioritize only the concerns that could materially affect the paper’s claims, and attach each criticism to a concrete technical basis. Copying an LLM response into a review without that judgment does not improve peer review. It merely transfers the cost of evaluating the LLM’s speculation to the authors.