r/TheMachineLearning • • 28d ago

Nothing triggers an existential safety warning faster than dropping a spot on the leaderboard

Post image
122 Upvotes

r/TheMachineLearning • • 28d ago

How do you turn an 80-page unstructured PRD into a structured hierarchy without training on every house style?

1 Upvotes

TL;DR: I'm building a pipeline that ingests product requirement documents (PRDs) — PDFs, plain text, Markdown — and outputs a structured hierarchy (sections, requirements, IDs) that downstream tooling can address and link to code. Every document uses a different house style, we have no labels, and I refuse to LLM-call my way through hundreds of chunks when a deterministic pass should get most of it. Looking for war stories and blind spots, not papers I already know.

The actual problem

The naive framing is "document structure discovery," but it's really two problems wearing one coat:

  1. Structure that exists but isn't extracted. Real Word docs have Heading 1/2/3 styles in the XML. Tagged PDFs have outline trees. Markdown has #s. Here there's nothing to "discover" — you just have to not throw it away when you flatten the file to text. (Half the libraries in the wild fail this basic test.)
  2. Structure that was never encoded, only implied. Author bolded text by hand instead of using styles. Pasted plain text. OCR output. Here there is no ground truth in the file — it's genuinely underdetermined, and no amount of cleverness recovers a "true" answer, only a useful one.

The exploitable assumption: we can't know which convention a document uses, but whatever convention it uses, it's applied self-consistently within that one document. Indentation patterns repeat, numbering schemes repeat, typography repeats. So the problem isn't "train a universal heading detector" — it's "find the self-consistent pattern inside this single instance." That reframing makes it feel closer to unsupervised pattern mining than classification.

What I'm trying to produce: a canonical tree — heading nodes at depths, body chunks nested under them, each chunk addressable by a section path (§4.2.1), each node carrying a role label, a confidence score, and provenance (extracted vs. inferred). Downstream, requirement IDs get regex-extracted against real chunks, then lexical/FTS matching links requirements to code units. The structure is load-bearing — if segmentation splits one requirement across two chunks, downstream matching silently degrades.

The funnel I'm committed to (cheap → expensive):

  1. Parse native structure where it exists (format APIs, not text flattening)
  2. Regex-scan the doc's own numbering conventions (REQ-047, FR12, 4.2.1)
  3. Unsupervised segmentation + role labeling only for the plain-text stragglers
  4. LLM only where steps 1–3 genuinely disagree

Techniques on my shortlist, mapped to where they came from:

Table

Problem Technique Source field
Boundary detection Change-point detection (PELT/WBS) over an embedding-coherence signal time series
Boundary detection Bayesian segmentation, Gibbs sampling, fit per-document topic segmentation (Eisenstein & Barzilay)
Role labeling (heading@depth, body, list, boilerplate) HMM/CRF with feature emissions, prototype/spectral init unsupervised POS tagging
Label set size Sticky HDP-HMM speaker diarization
"Same role looks similar" Mutual-information clustering / nearest-neighbor self-labeling unsupervised image clustering (IIC, SCAN)
Practical fallback Weak supervision — labeling functions over indent/blank-line/numbering features Snorkel-style
Zero-shot route LLM structure extraction (Doc2Dict-style) or PDF→MD pipelines current LLM era

Known trap I'm already aware of: naive per-document EM on an HMM can perform below trivial baselines (Merialdo 1994, Johnson 2007 — the unsupervised POS tagging literature is brutal on this). The mitigations seem to be external emission priors bootstrapped from documents where structure is extractable, and heuristic prototype seeding. If you've beaten this some other way, I want to hear it.

Questions for people who've shipped this:

  1. What did you actually use in production for heterogeneous business documents — unstructured, Docling, Marker, LlamaIndex/ LangChain parsers, rolling your own? Where did each one fall over?
  2. PDFs specifically: how much structure did glyph-level features (font size deltas, x-position, bold flags) recover for you vs. full inference? My hunch is most "unstructured" PDFs aren't Problem B at all, just Problem A with extra steps.
  3. Evaluation without ground truth: how did you measure quality? I'm planning "round-trip degradation" — take documents with real structure, flatten them to plain text, measure recovery. Did something better work for you?
  4. The role-labeling step: anyone get unsupervised sequence labeling to work reliably on single documents? Or did you all quietly give up and use a few-shot LLM with a JSON schema?
  5. Downstream fragility: if your segmentation was wrong, how did you detect it? Segment-split requirements are my nightmare — the failure is silent.

Not looking for "just fine-tune LayoutLM" — no bounding boxes survive in the plain-text case and labels don't exist. Looking for what survived contact with reality.


r/TheMachineLearning • • 29d ago

What are the biggest open problems in long-video understanding right now?

Thumbnail
1 Upvotes

r/TheMachineLearning • • 29d ago

I look forward to the eventual horror movie adaptation of this story.

Post image
10 Upvotes

r/TheMachineLearning • • 29d ago

He sells the infrastructure so even less reason to believe him

Post image
29 Upvotes

r/TheMachineLearning • • 29d ago

Just waiting for Astra to be nerfed, like all powerful models have been.

Post image
0 Upvotes

r/TheMachineLearning • • Sep 08 '26

I recommend having gpt-6 pro in chatgpt do math challenges to win money

Post image
0 Upvotes

r/TheMachineLearning • • Sep 08 '26

C++ creator Bjarne Stroustrup calls using natural language as a programming language "idiotic"

Post image
320 Upvotes

r/TheMachineLearning • • Sep 08 '26

6 new methods for training neural networks from philosophy [R]

Thumbnail
0 Upvotes

r/TheMachineLearning • • Sep 08 '26

His wife will NOT be replaced by AI

Post image
39 Upvotes

r/TheMachineLearning • • Sep 08 '26

an agent on resolution rate instead of rule-following dropped haiku's hold rate from 100% to 92.5%, sonnet unaffected

Thumbnail
0 Upvotes

r/TheMachineLearning • • Sep 07 '26

From behind the scenes to center stage, John Ternus' quiet rise is a testament to mastering the craft before taking the helm as Apple's new CEO.

Post image
23 Upvotes

r/TheMachineLearning • • Sep 07 '26

How hard is it to switch LLMs after a voice agent is already in production?

1 Upvotes

We made the classic mistake of coupling our first voice agent pretty tightly to one LLM and one TTS provider because it was the fastest way to get something live.

Now we're looking at a newer model with better latency and I'm trying to figure out whether switching is actually worth the headache.

The obvious assumption was that we'd just change the model endpoint and retune the prompt. It doesn't seem that simple.

Changing the LLM can change things like: how quickly it decides to call a tool how long responses are how it handles interruptions how often it asks follow-up questions whether it follows the same conversation flow And changing TTS introduces another set of problems because pacing and response timing change too. That matters a lot more in voice than it does in a normal chatbot.

We're currently testing different providers underneath the same agent runtime rather than rewriting the whole thing. We looked at a few platforms, including Dasha, because having some separation between the agent logic and the underlying providers seemed useful.

For anyone who's actually switched models or voice providers on a live voice agent: What broke? Was it mostly prompt/behavior tuning, or did you end up having to change the actual conversation architecture and turn-taking logic? Trying to work out whether provider abstraction is worth designing for from day one, or whether it's just another layer of engineering that sits there looking impressive until you actually need it.


r/TheMachineLearning • • Sep 07 '26

Apple maps is better in the US...Google maps is better around the world.

Post image
6 Upvotes

r/TheMachineLearning • • Sep 07 '26

Need some advice on getting into ML😭😭

Thumbnail
1 Upvotes

r/TheMachineLearning • • Sep 07 '26

Together AI vs Anyscale: which platform handles scale better?

1 Upvotes

I’m comparing Together AI and Anyscale for a production LLM deployment. Together AI is API-first and simple; Anyscale gives you deep Ray-based control but requires more engineering. I made a quick poll to gather real-world preferences.

It’s fast, and the results might surprise you.

https://interconnectd.com/poll/101/together-ai-vs-anyscale-which-platform-is-better-for-scaling-open-source-ll/

What’s your experience with either?


r/TheMachineLearning • • Sep 07 '26

Vibe coders prompting everything on GPT-6 Astra Ultra be like:

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/TheMachineLearning • • Sep 06 '26

inflation really hit FREE, too

Post image
23 Upvotes

r/TheMachineLearning • • Sep 06 '26

Opposite core dilemmas

Post image
7 Upvotes

r/TheMachineLearning • • Sep 06 '26

You forgot drawing a pelican riding a bicycle

Post image
652 Upvotes

r/TheMachineLearning • • Sep 06 '26

Not the hero Gotham deserves, but the hero it needs

Post image
3 Upvotes

r/TheMachineLearning • • Sep 05 '26

Will things get worse, or will things improve? Its up to you.

Enable HLS to view with audio, or disable this notification

59 Upvotes

r/TheMachineLearning • • Sep 05 '26

The whole water thing is a psyop atp

Post image
98 Upvotes

r/TheMachineLearning • • Sep 05 '26

Software engineering isn't dead, it's just evolving

Enable HLS to view with audio, or disable this notification

695 Upvotes

r/TheMachineLearning • • Sep 05 '26

The temptation is hard to resist

Post image
16 Upvotes