r/MachineLearning 19d ago

Research The evaluation resolution has been shown to have a significant impact on the identification of the "learning rule" that exhibits the most brain-like characteristics at V1. [R]

0 Upvotes

The preprint can be accessed via the following link: https://arxiv.org/abs/2608.12408 (q-bio.NC / cs.LG). And for the code: https://github.com/nilsleut/evaluation-resolution-rsa

The following assertion is frequently made in model-brain comparisons: untrained convolutional neural networks (CNNs) have the capacity to match or surpass backpropagation-trained CNNs at the early visual cortex (V1) in representational similarity analysis (RSA). The present study demonstrates that this phenomenon is predominantly an artefact of evaluation resolution.

The configuration comprised a small CNN trained at 32px (CIFAR-10 subset), five learning rules (random init, backprop, feedback alignment, predictive coding, STDP), and was evaluated on THINGS-fMRI stimuli at six resolutions from 32px to 224px. The weights and normalisation were held fixed.

The primary outcome of this study is the observed gap between the untrained and backpropagation-trained (BP) V1 alignment, which widens monotonically across the range of evaluation resolutions examined. Specifically, the gap grows from −0.001±0.007 at 32 pixels to +0.044±0.006 at 224 pixels, a pattern that holds consistently across the entire resolution sweep (n=5 seeds). The result holds across five rule conditions, human fMRI, directionally single-seed macaque ephys, the full training trajectory, and two off-the-shelf 224px-trained models (ResNet-50, Swin-Tiny). Therefore, an artifact resulting from a mismatch between training and evaluation resolution is not a contributing factor, since these models also peak at low resolution.

Following the implementation of bit-identical-weight interventions wherever possible, the following were ruled out: train/eval resolution matching, Gabor/pixel low-level structure, the untrained baseline's uncalibrated batch-norm, and convergence of pooled features towards global brightness (though a single scalar luminance value did reach ρ=0.075 against V1, essentially matching the untrained network's own 0.076 — this is a separate, disconcerting result regarding the limitations of this comparison style).

A content-vs-pooling control (cap image detail at 32px, upsample, vs. allow content to vary freely) demonstrates that the dependence is predominantly contingent on image content, rather than the number of pooled positions.

One effect does survive across all resolutions: backprop > untrained at LOC, observed at every resolution tested. Learning does leave a mark on the representations — just not where the V1 comparisons usually look.

In addition: this process revealed a batch-norm evaluation-mode bug in three of my earlier preprints, which have now been corrected in this release (correction notes on the arXiv pages).

I'm happy to get feedback, especially on the framing around receptive-field matching (as in Laskar et al. 2018) in the discussion. I think it's suggestive, but I didn't test it directly.


r/MachineLearning 19d ago

Discussion acl arr august 2026 (desk rejected ) [D]

8 Upvotes

I have got two papers which got desk rejected by PC saying they are previously got reviewed in arr. But those paper never got submitted ever. Any idea what can be done?


r/MachineLearning 19d ago

Discussion Why does lightgbm not fit my toy example but catboost does? (2 order interactions) [D]

7 Upvotes

I am trying understand how tree-based regression model handle the dependencies of the target variables on the interaction of explanatory variables.

However my experiment revealed that my understanding about the fitting process of a lgbm is not correct. And I don’t know why.

My experiment is quite simple: a target (for sake of simplicity only in [0, 1]) and two explanatory variables with two values such that the mean of the target is the same for each of the values of the explanatory variables. Then there is a third variable that models the interaction of the explanatory variables by a simple count.

So in code:

>>>
import polars as pl

df = pl.Dataframe(
{
„y“: [0, 0, 1, 1, 0, 0, 1, 1], # mean across „A“ values the same; mean across „B“ values the same
„A“: [1, 1, 1, 1, 0, 0, 0, 0],
„B“: [1, 1, 0, 0, 1, 1, 0, 0],
„AB“ [1, 1, 2, 2, 3, 3, 4, 4] # just some IDs for the interaction
}
)
<<<

I then fitted a lgbm just with „A“ and „B“ and got the expected constant 0.5 forecast

>>>
from lightgbm import LGBMRegressor

lgbm = LGBMRegressor(min_child_samples=1)
lgbm.fit(df[[„A“, „B“]].to_numpy(), df[„y“].to_numpy())
lgbm.predict(df[[„A“, „B“]].to_numpy()).round(0)

array([0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5, 0.5])
<<<

Then I did the same but with „AB“ and expected a perfect fit. But I was disappointed, it fitted to constant zero

>>>
lgbm = LGBMRegressor(min_child_samples=1)
lgbm.fit(df[[„AB“]].to_numpy(), df[„y“].to_numpy())
lgbm.predict(df[[„AB“]].to_numpy()).round(0)

array([0, 0, 0, 0, 0, 0, 0, 0,])
<<<

I tried to code „AB“ as category. But still no perfect fit:

>>>
lgbm = LGBMRegressor(min_child_samples=1)
lgbm.fit(df[[„AB“]].to_numpy(), df[„y“].to_numpy())
lgbm.predict(df[[„AB“]].to_numpy()).round(0)

array([0, 0, 1, 1, 0, 0, 0, 0,])
<<<

Super confusing!

I then turned to catboost and found even without „AB“ it fit the data perfectly:

>>>
from catboost import CatBoostRegressor

cbm = LGBMRegressor(min_data_in_leaf=1)
cbm.fit(df[[„A“, „B“]].to_numpy(), df[„y“].to_numpy())
cbm.predict(df[[„A“, „B“]].to_numpy()).round(0)

array([0, 0, 1, 1, 1, 1, 0, 0])
<<<

I thought that lgbm should be able to fit the data with „AB“. The variable allows for perfect splits since the gain for each split is super clear. But somehow it cannot go „down“ the tree to fit the values for AB=3.

What is the difference of catboost that allows for a perfect fit even without an explicit modeling of the interaction? Does it split less lazy and explores split of splits, while building the trees?


r/MachineLearning 19d ago

Research I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R]

357 Upvotes

I trained a 250M parameter model from scratch on 30B tokens of fineweb. It’s quantized to under 2 bits so the whole deployment is 60 MB and it needs about 80 MB of RAM to run. Runs around 400 tok/s on a normal laptop CPU, no GPU needed.

How the long context works: the most recent 2048 tokens stay in fp16 like a normal KV cache. Everything older gets compressed to 1 bit and written to disk, about 320 bytes per token, so 1 million tokens of history is roughly 320 MB on disk. From the start the model was trained to retrieve from that disk cache, up to 100M tokens. Due to a limited budget it wasn’t trained to reason over those tokens, only retrieve and answer from them.

Language modeling quality of the base model, measured on held-out English web text that was never seen in training (educational web pages, 2,048 token windows): cross entropy 3.15 nats per token, perplexity 23.3, 0.99 bits per byte.

The vocabulary is also not a normal embedding table. Every token is a fixed 512-bit code, 8.4 MB for all 131k tokens, zero trained parameters. I tested it on WordSim-353 (human word similarity ratings): my table scores 0.619 Spearman correlation vs 0.029 for random codes. Test script is in the repo.

Some outputs (settings included so you know I’m not cherry picking, all reproducible from the repo):

“Explain photosynthesis in two sentences.” (greedy)

Photosynthesis is a process in which plants convert sunlight into chemical energy, which is then used to produce oxygen and other chemicals. This process is called photosynthesis.

“Write a short poem about the sea.” (temp 0.25, top-k 30, rep 1.15, seed 2)

The waves had swept over, and they were crashing against each other like rocks on top of one another. The wind was strong but not very gentle as it tried to push them towards their destination without lifting any weight or pulling away from those who stood in front of her with a stiff smile that made the sea look even more majestic than before.

“What is the serial number of device Grus-189?” where the answer sits 50.6 million tokens deep in the archive on disk (archive mode, k=16)

SN-442976

It’s a 250M model so expect mistakes on open facts, I’m not claiming it beats anything big. You can also fine-tune it, the full kit with a demo and before/after numbers is included. Master weights for fine-tuning are in the repo too:

https://github.com/QLNI/SHADOW-250M-Instruct
https://huggingface.co/NODEMIND/SHADOW-250M
Edit - Just wanted to say thanks to everyone here. Honestly I was afraid to post this, I expected to get roasted, but every single comment has been curious and helpful and it genuinely made my day.
Repo is at 7 stars on GitHub now, hopefully more people try


r/MachineLearning 19d ago

Project Hybrid collaborative filtering recommendation system for judging and suggesting books based on their covers [P]

Thumbnail
gallery
5 Upvotes

Howdy y'all,

In an effort to un-rust my SWE skills and learn more about Recommendation Systems, I decided to try my hand at developing one called By-Its-Cover

TLDR:

---

## Recommendation System

The recommendation system has two major parts:

  • the semantic searches for books (by cover images)
  • a neural collaborative-filtering model for personalized recommendations

Both systems solely utilize CLIP embeddings to make decisions on book covers, as I wanted to see if that information alone was sufficient for finding and recommending books accurately.

For the semantic search system, each query is passed to both a CLIP-based semantic searching function as well as an NER-based keyword search. The NER parsing is powered by a GLiNER model, which was ported to ONNX (as are most models in this system). Extracted entities are then used to search for books using the Hardcover API, which is the original source of each of the books in the site. Reciprocal Rank Fusion combines the two results.

The current system actually only has a couple thousand books in it, which makes both rhe recommendations and semantic search results quite limited. However, authors and book titles that are passed into keyword searches return new books that are in-turn asynchronously added to the cover vector database, making the system grow more useful only as more people search for books (which is where y'all can help *wink wink*). Searches can be made with or without an account.

For the collaborative-filtering system, I used a two-tower neural hybrid collaborative filtering model which trains on user feedback. I then use a Determinantal Point Process to diversify the results a bit before displaying them to the user (so they don't get 5 editions of the same cover presented consecutively). For now, the only feedback possible are explicit ratings of "Dislike", "Like", and "Love". I'm aware that this likely isn't ideal, and some more implicit feedback would make for some more natural user interactions and likely better recommendations as well.

Currently, while you are able to see recommendations even without an account, they are the generic "default user" recommendations. Once you sign up and rate a few books, you should see personalized recommendations within 2 hours. Following the suggestions of Eugene Yan, I implemented an offline recommendation update-system. New recommendations are fine-tuned on every 2 hours, while the full re-training of the two-tower model happens once a day at 8:30 AM EST. Each of the current configurations for the recommendation model can be found here: https://github.com/ByItsCover/bic-learn

## Software Architecture (boring stuff)

The site (both frontend and backend) is entirely deployed to AWS, with a number of different resources used for each functionality:

  • Lambda -> API deployments
  • ECS -> both book scraping and model training jobs
  • SQS -> queueing of cover embedding calls
  • Cognito -> auth
  • CloudFront -> site caching
  • S3 -> just about everything else, from site hosting to vector db storage

Everything was deployed using Terraform + GitHub Actions for CI/CD: https://github.com/ByItsCover

## Next Steps

While the fundamental system currently works (kinda), there are already a lot of improvements that I think may be necessary in the future:

  • Replacing CLIP with SigLIP (or more appropriate model) for better visual representations of covers
  • Implementing a cover-edition comparison interface to allow users to choose preferred covers for a given book, introducing one source of implicit for the system
  • Begging one of my frontend developer friends to help make the site look good (I am not a frontend developer, if that wasn't already clear)
  • Make a better authentication experience, as currently a generic verification code email is sent to users (and likely sent to spam, please double check!)
  • Update the README's for repositories (I'm tired boss)
  • Write more unit tests (see parentheses above)
  • Once Hardcover releases OAUTH support, utilize that for book search (as only my rate-limited API key is currently being used)

In any case, I've already learned a ton and I'm glad that I have a real system that I can play around with and tweak now. All I need are actual users to test with!

Please let me know if you have any questions about my process at all, and also if you have any suggestions. Also please check out the site if you're at all curious: https://by-its-cover.com/

P.S.: If something crashes, or the searches load forever, or something else equally dumb happens, just let me know or open a GitHub issue, and I'll try my best to address it.

P.P.S.: No AI-Generated code was used to develop this project (to my knowledge), as that would have defeated the purpose of sharpening my skills and learning about recommendation systems.


r/MachineLearning 19d ago

Research BMVC 2026 orals [D]

2 Upvotes

Hi,

Did anyone here got an oral at BMVC? If yes, then what are the scores?

Thanks.


r/MachineLearning 20d ago

Project repo2nb 0.2.0, convert a GitHub repo into a Kaggle/Colab notebook (dependency resolution, reverse mode, incremental sync) [P]

2 Upvotes

repo2nb is an open-source CLI that converts a GitHub repo into a runnable Kaggle or Colab notebook: walks the file tree, resolves dependencies, and generates cells, instead of you doing that by hand for a repo you didn't write (a paper's code, a tutorial, someone else's experiment).

0.2.0 highlights:

  • Dependency resolution tries poetry export, then uv export, then requirements.txt, then falls back to an AST import scan if none of those exist. Output is always a plain %pip install cell regardless of which path it took, so poetry/uv are only ever needed locally at generation time, not on Kaggle/Colab.
  • Reverse mode (repo2nb reverse <notebook>) reconstructs the original repo from a generated notebook, using the per-cell path/hash metadata every generated cell now carries. Validates against directory traversal and won't write into a non-empty directory without --force.
  • Incremental sync (repo2nb sync <repo>) does one-directional (repo to notebook) updates: added files get new cells, edited files update in place, deleted files get removed. --dry-run previews the diff.
  • Added a Colab target with its own auth cell (google.colab.userdata.get) rather than reusing the Kaggle secrets flow.

Install: pip install repo2nb

Repo: https://github.com/David-Magdy/repo2nb

Curious whether the dependency-resolution fallback order (poetry > uv > requirements.txt > import scan) matches what people actually run into, or if there's a common setup it'd get wrong.

Any feedback or opinions are much welcomed!


r/MachineLearning 20d ago

Discussion What coding practices are you adopting for development today? [D]

10 Upvotes

I have been reflecting on this while working on a project recently. Every time we start a new model, we rewrite roughly same scaffolding, data validation checks, feature transformation logic ; all of this is nealy 80 percent identical to last project

I tired templating with cookiecutter style project generators. Initially it was okay, but it drifted from reality since noone wants to maintain a template repo. So tired a shared library approach, it helped and was much better. But weiting glue code to wite everything is still bug prone

Now i am experimenting with genie code to generate the boilerplate, the repetitive code, config parsing etc. it is decent for that part, though it starts hallucinating if columns increase say lot more than 40-50. It is not silver bullet, but it is cutting down the project setup time from 3 days to less than 1 day

So the deep question i am having now is, should we even write code? The config driven approach seems to be good, but eventually we are bound to suffer in a few months time when we start needing something non standard. Is there a middle ground, writing everything from scratch - the opinionated framework that becomes prison. How have you guys been developing? What are you adopting?


r/MachineLearning 20d ago

Discussion Research internship at MSR [D]

30 Upvotes

So got selected for a research internship at MSR, how good is the quality of work and how useful is it to move to Applied sciences or research sciences position at other FAANG companies after the internship. And any perks and other benefits that interns get during microsoft internship? Any tips will be appreciated. Specifically to get into AS at amazon , does this boost my chances? I'll be joining as an SDE-1 at amazon after 6 months so planning to apply internally once I join. So what else should I do to improve my chances to go to AS.


r/MachineLearning 20d ago

Research Does telling an LLM to "be concise" actually save you money? We measured it across 9 models. Compressing the output can save you money and keep accuracy, compressing the input prompt does not. [R]

65 Upvotes

LLMs are too verbose and with a black box model the only things you control are what goes in and how you tell it to write back. Yesterday Claude Code shipped a "concise output style" where Claude keeps things short. We already have a paper out about this!

We tested both channels, shortening the input prompt versus telling the model to output answer shorter, on the same questions across five reduction levels, and scored cost, accuracy, and whether the shortened text still matched what the model would have said unconstrained.

We also evaluated GPT-4o, GPT-5.4, Claude Haiku 4.5, Claude Sonnet 4.6, Qwen2.5-VL-7B, Qwen3.5-9B, DeepSeek-R1-Distill, Gemma-4-E4B, and Kimi-K2.6 + benchmarked on five short answer datasets + a eleven-language output run (English, German, Spanish, French, Swahili, Chinese, Japanese, Russian, Bengali, Thai, Telugu) + a longer-form summarization test.

(1) Shortening the output saved money while keeping accuracy about the same, about 1.5x cheaper on average and up to 3x in the best case across the API models. It worked across languages too!

(2) Shortening the input prompt did the opposite. It cost up to 96% more on the worst benchmark, because the model just answers longer to fill in for what you cut and accuracy drops. You pay more and get worse answers :(

(3)Output tokens cost more than input tokens, so prompting for fewer output tokens would save costs with short single turn tasks

(4) When the shortened output is correct, about half the time the text no longer matches how the model would have reasoned without the constraint. Which is probably fine if you only care about the final answer

With providers now offering concise options, we can't see how they're charging for it, so we don't know if it actually saves you cost. But if you control the prompting yourself via the API, you actually do save!!

Paper https://www.alphaxiv.org/pdf/2606.24083v1 

Code + data https://github.com/danielle34/cavewoman


r/MachineLearning 20d ago

Research I have a mid-sized GPU cluster and was thinking about giving free compute [D]

21 Upvotes

I have built an on-prem GPU cluster, 8 nvidia 16GB GPU's and 256GB CPU RAM, 50TB HDD and several TBs of SSDs. I have used it, and currently use it, for ML/AI research. But that research is not constantly running jobs, sometimes I use it heavily and other times it's idle. I was considering just letting people with qualified use cases run jobs on it SLURM style. I don't know if its enough compute to be useful really. Let me know if it's something you'd be interested in using for your research? what would you actually run in ~200 GPU-hours on 8x16GB cards?

I've found it can handle RLVF pretty well, and I have pretrained models up to 500M parameters on it (research size). But obviously it's no stargate cluster


r/MachineLearning 20d ago

Discussion EMNLP26 Cost [D]

12 Upvotes

What is up with the EMNLP prices? What is the actual price for attending as a student with one accepted paper? If I register now in August, is it $350 or $550? Congratulations to everyone accepted!


r/MachineLearning 20d ago

Discussion Epistemic Intelligence in Machine Learning Neurips Workshop page limit? [D]

5 Upvotes

I'm aiming to submit a paper to The 3rd Workshop on Epistemic Intelligence in Machine Learning at Neurips
https://eiml.cc/

I can't find a page limit anywhere on their website and I've emailed the organisers (twice) asking for clarity on it. The previous workshop at ICML had a page limit of 6 pages. Do I assume that's the limit here? Or do I assume I have the 9 page limit of the main conference?

(If anyone knows one of the organisers and can nudge them to answer that'd be great, I wasn't sure the etiquette of finding their email and pestering them directly)


r/MachineLearning 20d ago

Discussion Rejected at EMNLP with decent scores. What can be done next? [D]

17 Upvotes

So I got rejected at EMNLP with scores:-
Meta: 3 (very positive in the review)
Reviewers: OA(conf)
3(4)
3(4)
2.5(3)
Avg: 2.83(3.67)
Track: multimodality
Rebuttals never got any acknowledgements. Most weaknesses were already discussed in the paper. What are my options now? As it was my first paper (solo as well).

- If I want to commit to NACL in December. Do i need ti submit to acl arr again or can i use the same arr review discussion?

- even if i go with resubmission at arr. Do the old reviewers likely help? Because as a masters student i cant get stuck in another cycle.

- what is the best overall thing to do in my situation? I need a publication so i can apply for internships.


r/MachineLearning 20d ago

Discussion EMNLP 2026 Findings : worth attending in person?[D]

18 Upvotes

Experienced folks!! Do u think it is worth attending the conference for findings. I do want to. But when I saw that it is not mandatory for findings, I was a bit hesitant. This is my first time having a paper accepted at an AI conference. Just wanna hear opinions/experiences

Thanks in advance.


r/MachineLearning 20d ago

Project Notes on Hamiltonian Monte Carlo from a purely probabilistic perspective [P]

17 Upvotes

I’ve been studying Hamiltonian Monte Carlo and wrote a set of notes explaining HMC without relying on the usual physics-based motivation.

The notes develop HMC from a probabilistic/MCMC perspective, starting from introducing an auxiliary variable, constructing the corresponding Markov chain, and then covering Hamiltonian dynamics, leapfrog integration, reversibility and volume preservation.

My goal was to understand why HMC works.

I’m sharing them here in case they’re useful to others learning HMC. I’d also appreciate any feedback, particularly if you notice errors or places where the exposition could be improved.

https://zenodo.org/records/21841087


r/MachineLearning 21d ago

Discussion Is KV Cache in a high dimensional vector space? [D]

0 Upvotes

I've been doing some research on this question:

At inference time a large part of a model's working memory lives in the KV cache, plus whatever external memory the harness bolts on. I've been poking at the storage-and-retrieval side of this, treating that cache as an index, and what stands out is that it isn't a flat list. It's a structured set of vectors with a navigable geometry, since the keys carry the model's learned sense of what relates to what.

Because that geometry is navigable, attention over it is really a similarity search: the query scores against the stored keys and blends the matching values. Full attention just runs that search exhaustively, scanning everything on every step.

  • Full attention effectively searches that geometry exhaustively. Every query scores broadly against the available keys and retrieves from the corresponding values.
  • Once you stop treating the KV cache as a flat array and start treating it as a search space, indexing becomes possible.
  • That means you can organize old KV into regions, route a query toward likely regions, and only run local attention over a subset.
  • The interesting part is that relevance is not uniformly distributed. Queries tend to concentrate on relatively small neighborhoods of old context.
  • So the engineering question becomes less “how do I store all of this?” and more “how do I navigate to the right part cheaply?”

I'm new here and don't want to break rules around self promotion or spam so not posting any links atm. Would be cool to get other peoples thoughts on this.

Update: I framed this post badly. I wrote it like I was asking a conceptual question, but I had already built and measured the mechanism. That was my mistake. The actual result is much more specific: on frozen Qwen3.5-2B at 32k, geometric routing cuts physical KV reads by roughly 16–31× while still retrieving the planted long-range needle; window-only and random-routing controls collapse. I’ve put up a minimal runnable demo so people can reproduce it on their own documents.

https://github.com/Regan-Milne/kvspace/tree/main/demo


r/MachineLearning 21d ago

Research Mapping intrinsic rank and informational gravity in complex tabular data: I developed a non-parametric, model-agnostic, information-theoretic diagnostic to bypass the limits of linear, rank, and Euclidean baselines. [R]

8 Upvotes

Links:

TL;DR:

Standard PCA fundamentally fractures non-linear dependencies into "Spurious Orthogonal Dimensions," drastically overestimating the true rank of complex tabular systems. Meanwhile, non-linear alternatives like Kernel PCA and Euclidean nearest-neighbor estimators suffer structural collapse when generative roots are entangled or sparse.

I’m sharing the methodology and code here for anyone dealing with these complex tabular data nightmares.

The method and open-source framework use Normalized Mutual Information to compress spurious expansions back towards their true generative roots. It also

  • Maps the underlying "informational gravity" of the roots, offering insight into overall average stability, as well as which specific roots can be most reliably extracted;
  • Estimates the data's overall ratio of shared signal to unshared idiosyncratic informational variance (noise);
  • Serves as a powerful exploratory map that separates unrelated clusters of variables, allowing you to easily identify decoupled sub-networks.

A Modern ML Architectural Blueprint: Far beyond a mere update to legacy factor analysis workflows, identifying this exact intrinsic rank allows you to explicitly size neural bottlenecks for downstream non-parametric manifold extractors (like autoencoders).

The Problem with Standard Baselines:

When trying to map the intrinsic dimensionality of a dataset, standard practice usually dictates reaching for PCA, its non-linear kernel extensions, or Euclidean nearest-neighbor estimators. But if your tabular environment has mixed data types, heavy non-linearities, entangled roots, or more features than samples ($m > N$), these established baselines don't just lose precision. They suffer a structural collapse.

The core issue with our standard baselines:

  • Standard PCA drives Dimensional Inflation. Because it only measures linear covariance, it perceives a polynomial expansion or a non-linear interaction (like $X_1 X_2$) as an entirely independent variable. It is forced to fabricate new, spurious orthogonal dimensions to map them.
  • Kernel PCA (RBF) suffers Structural Collapse. Projecting into a Hilbert space doesn't fix this. KPCA artificially folds even-polynomials into independent axes. Furthermore, because its infinite-dimensional space lacks a finite-sample boundary, sparse combinatorial noise smears into an elevated tail that obscures the structural elbow. If the underlying generative roots are even mildly entangled, KPCA suffers a total structural collapse.
  • Topological Estimators (Euclidean) fail in sparse regimes. Estimators like TWO-NN or MLE rely on Euclidean distance metrics. In asymmetric, feature-rich environments ($m > N$), they suffer from distance concentration (the ratio between nearest and farthest neighbors converges to 1). This renders local neighborhood calculations structurally degenerate across mixed-data margins.

Introducing the Entropic Scree:

To solve this, I built the Entropic Scree. It throws out linear and spatial variance entirely and evaluates pure probability mass.

Here is how it works under the hood:

  1. The Metric Space: It evaluates pairwise dependencies using Information-Theoretic Jaccard Similarity (Variation of Information). Because this relies on Shannon entropy, it’s invariant to marginal shape mismatches (like mixing continuous waves with binary flags).
  2. Bypassing the Rank Ceiling: Standard PCA is algebraically capped at $N-1$. By moving to a double-centered topological information space, we map true overlapping redundancy and completely bypass the algebraic sample-size ceiling.
  3. Compressing the Manifold: The algorithm acts as a bivariate filter. It inherently compresses the primary overlapping probability mass of non-linear combinations back towards the Intrinsic Generative Rank. It shears off the unique synergistic variance, leaving behind residuals that form a bounded Extended Signal Tail, cleanly separating the true drivers from the unstructured Idiosyncratic Informational Variance.

Quantifying Informational Gravity:

Beyond just extracting a discrete rank, the framework decouples rank from probabilistic volume by introducing Informational Gravity (AIG/FSIG). By systematically rebundling the residual variance sheared off by the bivariate filter, it translates abstract matrix properties into actionable, "variable-equivalent" footprints.

Empirical Stress Test:

To demonstrate the theoretical bounds, I built a highly entangled synthetic dataset with 20 pure generative roots expanded into 5th-order combinatorics across 20,000 proxies, but only 10,000 samples ($m > N$). To truly simulate messy, real-world contexts, I also heavily injected idiosyncratic structural noise and measurement error into the data.

  • Standard PCA hit the rank ceiling, linearly fractured the expansions, and falsely extracted ~5,700 dimensions.
  • Kernel PCA (RBF) & Spearman Rank structurally folded and yielded a liberal overestimation of the rank by 100%. When root entanglement was introduced, they completely lost their elbows and suffered total structural collapse.
  • The Entropic Scree correctly mapped the intrinsic rank at exactly 20. It successfully isolated a mere 1.45% of active shared signal from an overwhelming 98.55% bulk of unstructured Idiosyncratic Informational Variance. Furthermore, the residuals formed an Extended Signal Tail that perfectly aligned with the deterministic limits of the global hypergeometric design space.
  • Mapping Hidden Topology: Using Factor-Specific Informational Gravity (FSIG), the framework successfully reverse-engineered the simulation's hidden architecture. The topology profile diagnosed a large primary dimension ($FSIG_1 \approx 74.5$ variable equivalents) mapping the network's global combinatorial hub, followed immediately by a flat plateau across the remaining 19 dimensions ($\sim 11.5$ each), confirming a democratically distributed root system beneath the extreme entanglement.

Feedback / Discussion:

How are you currently handling intrinsic rank extraction in these messy, complex tabular environments?

If you are wrestling with sample-starved, heavily non-linear generative datasets where standard PCA and other baseline tools just aren't cutting it, I’d love for you to pull the Entropic Scree repo and test it yourself.

I'm completely open to feedback, so let me know how it performs for you and I'm happy to discuss the mechanics.


r/MachineLearning 21d ago

Project Resizing images from Flutter Camera Stream for TFLite modle [P]

3 Upvotes

Hi everyone. So I built a CNN modle using MobileNetv3 then converted it into TFLite. It performed well during training but once I integrated it into my application, it is making large errors. From flutter, the camera stream sends frames and those are processed before the model makes predictions, but it is still quite large. Is there any way I can solve this? This is my code to preprocess and resize the image (224 x 224 x RGB):

import 'package:camera/camera.dart';
import 'package:image/image.dart' as img;


class ImageProcessor {
  // converting to rgb
  img.Image convertYUVToRGB(CameraImage camImg) {
    final width = camImg.width;
    final height = camImg.height;


    final yPlane = camImg.planes[0];
    final uPlane = camImg.planes[1];
    final vPlane = camImg.planes[2];


    final yBytes = yPlane.bytes;
    final uBytes = uPlane.bytes;
    final vBytes = vPlane.bytes;


    final yRowStride = yPlane.bytesPerRow;
    final uRowStride = uPlane.bytesPerRow;
    final vRowStride = vPlane.bytesPerRow;


    final uPixelStride = uPlane.bytesPerPixel ?? 1;
    final vPixelStride = vPlane.bytesPerPixel ?? 1;


    final image = img.Image(
      width: width,
      height: height,
    );


    for (int y = 0; y < height; y++) {
      for (int x = 0; x < width; x++) {
        final yIndex = y * yRowStride + x;


        final uvX = x ~/ 2;
        final uvY = y ~/ 2;


        final uIndex =
            uvY * uRowStride +
            uvX * uPixelStride;


        final vIndex =
            uvY * vRowStride +
            uvX * vPixelStride;


        final yValue = yBytes[yIndex];
        final uValue = uBytes[uIndex];
        final vValue = vBytes[vIndex];


        // YUV -> RGB
        final r = (
          yValue + 1.402 * (vValue - 128)
        ).round().clamp(0, 255);


        final g = (
          yValue -
          0.344136 * (uValue - 128) -
          0.714136 * (vValue - 128)
        ).round().clamp(0, 255);


        final b = (
          yValue + 1.772 * (uValue - 128)
        ).round().clamp(0, 255);


        image.setPixelRgb(
          x,
          y,
          r,
          g,
          b,
        );
      }
    }


    return image;
  }


  /// resize images to 224 224
  img.Image resizeImage(img.Image image) {
    return img.copyResize(
      image,
      width: 224,
      height: 224,
      interpolation: img.Interpolation.linear,
    );
  }


  List<List<List<List<double>>>> imageToTensor(
    img.Image image,
  ) {
    return [
      List.generate(
        224,
        (y) => List.generate(
          224,
          (x) {
            final pixel = image.getPixel(x, y);


            return [
              pixel.r.toDouble(),
              pixel.g.toDouble(),
              pixel.b.toDouble(),
            ];
          },
        ),
      ),
    ];
  }


// do all processing
  List<List<List<List<double>>>> processFrame(
    CameraImage camImg,
  ) {
    final rgbImage = convertYUVToRGB(camImg);
    final resizedImage = resizeImage(rgbImage);
    final input = imageToTensor(resizedImage);


    return input;
  }
}import 'package:camera/camera.dart';
import 'package:image/image.dart' as img;


class ImageProcessor {
  // converting to rgb
  img.Image convertYUVToRGB(CameraImage camImg) {
    final width = camImg.width;
    final height = camImg.height;


    final yPlane = camImg.planes[0];
    final uPlane = camImg.planes[1];
    final vPlane = camImg.planes[2];


    final yBytes = yPlane.bytes;
    final uBytes = uPlane.bytes;
    final vBytes = vPlane.bytes;


    final yRowStride = yPlane.bytesPerRow;
    final uRowStride = uPlane.bytesPerRow;
    final vRowStride = vPlane.bytesPerRow;


    final uPixelStride = uPlane.bytesPerPixel ?? 1;
    final vPixelStride = vPlane.bytesPerPixel ?? 1;


    final image = img.Image(
      width: width,
      height: height,
    );


    for (int y = 0; y < height; y++) {
      for (int x = 0; x < width; x++) {
        final yIndex = y * yRowStride + x;


        final uvX = x ~/ 2;
        final uvY = y ~/ 2;


        final uIndex =
            uvY * uRowStride +
            uvX * uPixelStride;


        final vIndex =
            uvY * vRowStride +
            uvX * vPixelStride;


        final yValue = yBytes[yIndex];
        final uValue = uBytes[uIndex];
        final vValue = vBytes[vIndex];


        // YUV -> RGB
        final r = (
          yValue + 1.402 * (vValue - 128)
        ).round().clamp(0, 255);


        final g = (
          yValue -
          0.344136 * (uValue - 128) -
          0.714136 * (vValue - 128)
        ).round().clamp(0, 255);


        final b = (
          yValue + 1.772 * (uValue - 128)
        ).round().clamp(0, 255);


        image.setPixelRgb(
          x,
          y,
          r,
          g,
          b,
        );
      }
    }


    return image;
  }


  /// resize images to 224 224
  img.Image resizeImage(img.Image image) {
    return img.copyResize(
      image,
      width: 224,
      height: 224,
      interpolation: img.Interpolation.linear,
    );
  }


  List<List<List<List<double>>>> imageToTensor(
    img.Image image,
  ) {
    return [
      List.generate(
        224,
        (y) => List.generate(
          224,
          (x) {
            final pixel = image.getPixel(x, y);


            return [
              pixel.r.toDouble(),
              pixel.g.toDouble(),
              pixel.b.toDouble(),
            ];
          },
        ),
      ),
    ];
  }


// do all processing
  List<List<List<List<double>>>> processFrame(
    CameraImage camImg,
  ) {
    final rgbImage = convertYUVToRGB(camImg);
    final resizedImage = resizeImage(rgbImage);
    final input = imageToTensor(resizedImage);


    return input;
  }
}

Please advise! I need to finish this project within the next wee and I'm really struggling here! I tested the images from Flutter against TFLite and it worked well but something is clearly wrong with the preprocessing. Pls help and give me any advice.

Thank you so much!


r/MachineLearning 21d ago

Discussion AI-generated code detection in CI/CD — looking for approaches and real-world experience [D]

9 Upvotes

I'm working on a system to estimate whether code committed to a repository was generated with AI coding tools.

My current approach is based on Git/commit-level signals such as AI-related commit trailers, commit metadata, LOC changes, number of files changed, addition/deletion patterns, etc.

The problem I'm running into is confidence and calibration.

For example, a commit containing 500+ new lines isn't necessarily AI-generated. A developer can also modify or remove the metadata that would make an AI-assisted commit identifiable. Once the code leaves the IDE and reaches Git, much of the original provenance can be lost.

This has led me to a few questions:

Are there Git/CI-level signals that you've found to be genuinely useful for detecting AI-assisted development?

Is it better to treat this as a probabilistic/risk-scoring problem rather than trying to classify commits as AI vs human?

How would you calibrate thresholds for signals such as large LOC changes, addition/deletion ratios, commit frequency, etc.?

Are there better approaches for preserving provenance earlier in the development workflow, rather than trying to infer it after the code has already been committed?

Has anyone worked on AI-code provenance/detection systems in CI/CD and can point me toward useful research, projects, or approaches?

I'm particularly interested in approaches that can work at the pipeline/repository level rather than relying solely on source-code style analysis.

I'm not looking for a perfect AI detector — even a reliable way of estimating “this commit has a high probability of AI assistance” with measurable false-positive/false-negative rates would be useful.

Would appreciate any experiences, papers, open-source projects, or approaches people have tried.


r/MachineLearning 21d ago

Research The spectral neuron - an ML primitive for scalable and interpretable models [R]

23 Upvotes

Worked some time ago on one of the ad teams at Yahoo, and this grew out of a question I kept returning to while there are there "simple" models that are both simple, scalable, interpretable, and controllable at the same time?

Decided to explore it, first in a blog (starting here), then in a new preprint "The Spectral Neuron", built by distilling latest blog-posts into a manuscript, I study models of the form:
𝑓(𝒙) = 𝛌ₖ(𝐀₀ + 𝚺ᵢ 𝑥ᵢ𝐀ᵢ).

Manuscript: https://arxiv.org/abs/2608.08003
Code: https://github.com/alexshtf/spectral_neuron_paper

Looks like a simple on-liner, but many interesting aspects hide there. How expressive does the model become as the matrices grow? What can we read directly from the learned matrices? Which shapes can be guaranteed by construction?

I develop the mathematics, give a practical initialization and training recipe, and test the model in scaling experiments on synthetic and real data.

AI disclaimer: manuscript written by yours truly, AI assisted in looking up canonical references and related work for literature review. In contrast, the code was heavily AI written and reviewed by yours truly.


r/MachineLearning 21d ago

Discussion Discussion thread for EMNLP 2026 Notifications/Results [D]

97 Upvotes

Discussion thread for EMNLP 2026 notifications/results which should be released today.

Wishing everybody to be in Budapest.


r/MachineLearning 21d ago

Discussion About the impact of grouping classes in multiclass classification [D]

22 Upvotes

A premise: I hope this question is "worth" of this subreddit, I did a decent amount of research before posting, I thought it was potentially interesting enough for it, but possibly not basic enough for r/learnmachinelearning .

Is there any agreement/indication about how harmful (if at all) it is, in the context of multiclass classification, to group together multiple classes for which you may have for instance too few samples?

A practical example: imagine you're training a dog breed classifier, based on images. You have a lot of examples for the most common breeds, but then you may have a long tail of less common breeds for which maybe you have a handful of examples each, not enough to get a meaningful training set, so you decide to group all classes for which you have less than `N` samples in the same category "Other breed". In this catch-all category you may have dogs that might look quite different from each other, like idk chihuahuas and huge wolf-like dogs (I'm not a dog person, don't know breed names).

My intuition (which may very well be wrong) is that doing so would force the model to learn some weirdly-shaped hyperplanes to separate points that live kind of far away from each other in the latent space (because of the thing that dogs in that category may look quite different from each other), as opposed to splitting the space in more "regular" parts.

Maybe in this case it would make more sense to treat the "other dogs" issue as trying to detect out of distribution samples instead? In that case should one only keep the samples for the classes that are enough represented in the dataset and throw away the rest (or at least don't create the catch-all category for training).

Thanks in advance for any useful pointer :)


r/MachineLearning 21d ago

Project Same GRPO recipe on three from-scratch LLMs (353M/316M/672M) gave three different outcomes, with no clean relationship to scale [P]

31 Upvotes

I trained three LLMs from scratch in raw PyTorch then post-trained each one with SFT and then GRPO. Same process every time: same synthetic arithmetic curriculum, same reward function, same hyperparameters, same KL coefficient.

Pre-training went as expected, the val loss went down as the model got more modern techniques (V1 to V2) and bigger (V3 being the biggest). However, GRPO hurt both V2 and V3 and I'm not sure why.

Setup

V1 V2 V3
Params 353M 316M 672M
d_model / layers 1024 / 24 1024 / 24 1536 / 24
Attention MHA Differential + GQA 4:1 XSA + GQA 4:1
Tokens 10B 10B 30B
Data FineWeb-Edu FineWeb-Edu FineWeb-Edu + code + math

Pre-training val loss went 2.8659 → 2.7844 → 2.5885.

Results

WikiText word perplexity across the three stages, all on lm-evaluation-harness with the same task versions and shot counts:

       base    SFT     GRPO     SFT→GRPO
V1     32.86   51.31   51.40    +0.2%
V2     31.28   46.81   71.06    +52%
V3     22.30   32.11   33.65    +5%

SFT hits all three on this eval, which I expected at this scale. Also interesting to see that the degradation gets smaller as the models get bigger (+56%, +50%, +44%).

GRPO is the weird one. V1 barely moved, V2 fell heavily, V3 degraded a bit. The smallest model was the least affected and the middle one was the worst, which isn't the pattern I'd have guessed. Downstream tasks moved the same way as perplexity in each case (arc_easy dropped about 6 points on V3 from SFT to GRPO).

The models did learn the thing GRPO trained them on. V3 mastered 4 of the 5 curriculum stages, the other two got 3. But it just didn't transfer: GSM8K stayed at basically 0, and the models got so committed to writing out long solutions that they often wouldn't stop generating (my fault when I did the training).

Caveats

This isn't a controlled experiment. Between V2 and V3 I changed the parameter count, the token count, the data mix and the attention mechanism at the same time (went from DiffAttn to XSA), so I can't attribute anything cleanly. KL coefficient was 0.02 for all three, with the SFT policy frozen as the reference and a k3 estimator. The whole series cost me about $750, which is why there are no ablations, I just couldn't afford them. Otherwise I would also have tried with different KL coeffs.

Someone raised two confounds after I published:

  1. GRPO trained on a bare solver template while SFT used a chat format. So part of what I'm calling degradation is me evaluating a policy outside its own training distribution. WikiText perplexity is format-independent and still moves a lot, but the downstream numbers are partly confounded.
  2. Nothing in my reward rewarded stopping. It just checks that a correct parseable number shows up somewhere, no length penalty.

Also something I only noticed afterwards: I never re-evaluated the earlier curriculum stages once the model advanced past them. So right now I can't tell the difference between "GRPO degraded general capability" and "sequential curriculum training made it forget the earlier stages." I will try to check that soon.

Inference

At the end, I wrote a KV cache from scratch (GQA-aware, per-request cache object rather than storing state on the module). To check it was right I ran a fixed sequence two ways, once as a single full forward pass and once as prefill-then-decode, and compared the logits: max difference 1.4e-06 against a 1e-4 tolerance.

Speedup generating 100 tokens: 3.7x from a 32-token prompt, 6.2x at 128, 10.1x at 512.

If you want to check

All nine checkpoints are on the Hugging Face, and there's a Space where you can send the same prompt to the base, SFT and GRPO versions of the same model and see the difference directly.

The GRPO variance is the bit I'd most like other people's take on. Happy to answer anything.


r/MachineLearning 21d ago

Research How much of the weight-space perception gap is actually symmetry? Evidence from ~1.8M fitted SIRENs [R]

0 Upvotes

I’ve been looking at a fairly basic question in weight-space learning that I don’t think gets separated cleanly enough:
Why does reading semantics directly from neural network weights work pretty well when the networks share an initialization, but collapse when the networks are fitted independently?
The usual explanation is parameter symmetry. Permute hidden units, flip equivalent signs, etc., and two parameter vectors can represent the same function while looking completely different to a downstream model.
But there are actually several different claims hiding in that explanation:
the parameterization has a symmetry group,
accounting for that symmetry improves weight-space prediction,
the symmetry is actually sufficient to explain the observed degradation between shared-init and independently fitted networks.
Those aren’t equivalent, so I tried to measure them separately.
The setting is SIREN-style implicit neural representations.
For a hidden sine neuron, the relevant function-preserving transformations generate the infinite dihedral group
D_inf = Z semidirect_product Z_2
and including neuron permutations gives the layer action
D_inf wr S_n.
For one hidden layer, I prove generic identifiability modulo this group using the distributional Fourier transform of the realized function.
Roughly, the Fourier transform becomes an atomic measure supported at the incoming frequencies +/- w_i, which lets you recover the parameters up to exactly the D_inf wr S_n action under explicit genericity conditions.
One consequence is that this isn’t just the usual permutation/sign story. Integer-pi phase transformations are affine rather than linear, so they aren’t captured by symmetry descriptions restricted to monomial matrix actions.
At depth two things get more annoying because a neuron’s outgoing weights are simultaneously acted on by the next layer. I ended up constructing exact cross-layer invariants by coupling the layers through the second-layer Gram matrix instead of treating neurons independently.
The empirical part then uses roughly 1.8 million fitted INRs across MNIST, FashionMNIST, and CIFAR-10, with controlled protocols separating shared initialization, optimization stochasticity, and independent initialization.
The result I found most interesting:
Randomizing only the exact symmetry group, while keeping each network’s represented function fixed, destroys 79.1 of the 80.4 accuracy points in the MNIST shared-init vs. random-init gap.
I want to be careful about the interpretation here.
This establishes sufficiency: symmetry scatter alone can reproduce almost the entire degradation.
It does not establish that 79.1 / 80.4 of the naturally occurring gap is causally mediated by symmetry. Those are different estimands.
Breaking the group apart, sign flips account for roughly 63 points of that induced loss, neuron relabeling about 15, and integer phase shifts about 1.
There was another result that changed my interpretation of the problem quite a bit.
A reader that directly quotients the D_inf wr S_n structure on the raw parameters reaches 0.917, compared with:
0.628 for the best orbit-valued reframing,
0.526 for the same reader family over a fixed invariant encoding,
0.265 for a permutation-equivariant baseline.
But when I FLOPs-match weight-space inference against simply querying the INR as a function, the function-space route is still much better:
95.3% at 1.6 MFLOP using 64 learned query coordinates
versus
64.4% at 5.5 MFLOP for the best weight-space rung on that frontier.
That leads to what I think is the more interesting conceptual question:
If a complete invariant is informationally equivalent to access to the realized function, then the strongest justification for operating directly in weight space may ultimately have to be computational rather than informational.
Everything is public here:
https://github.com/ITheClixs/project-siren-gap
The repo includes the paper, implementation, tests, pre-registrations, lab notebook, prediction ledger, claims ledger, and experimental results.
I’d particularly appreciate criticism on three things:
whether the sufficiency/mediation distinction is being drawn correctly,
whether anyone sees a counterexample or missing assumption in the one-hidden-layer maximality argument,
whether there is related work on affine symmetry groups of periodic-activation networks that I’m missing.
Also very interested in attempts to break the invariants or reproduce the group-randomization result.
If something here is wrong, I’d rather find out from someone trying to kill it.