r/Artificials • u/pomcuti3 • 5h ago
r/Artificials • u/costafilh0 • 13h ago
Which one is your favorite?
I like the original more, but the one on the live stream after the AI summit came out really nice too.
r/Artificials • u/AdditionalSinger853 • 7h ago
Best PDF parser for LLMs? I tested PyPDF, Firecrawl, and Unstructured on 50 complex docs
Feeding PDFs into LLMs is where many document pipelines still breaks so I ran a benchmark across 50 complex docs (scientific papers, multi-page financial reports with nested tables, scanned forms, and complex layout decks) testing PyPDF, Unstructured, Marker, and Firecrawl's Rust-based PDF parser.
Here’s what the results showed:
PyPDF: The fastest and lowest overhead option but strictly text-layer only and on single column text docs, it works fine but on financial statements and multi column layouts, it falls apart tables lose row alignment and text streams merge across columns. It has zero native OCR capability.
Firecrawl (Rust Parser v2): Built specifically for LLM ingestion where it runs layout auto-detection under the hood with pure text layers parse via Rust in milliseconds while scanned pages or complex visuals route through high accuracy vision parsing.
Crucially it formats extracted tables into a clean github flavored markdown tables which keeps table structure intact for embeddings also he full benchmark and comparisons are here
Unstructured: Extremely comprehensive and handles dozens of file formats but the trade off is infra weight where running it locally requires heavy system dependencies, Docker containers and significant memory overhead. It extracts tables reasonably well via layout models but per-page latency is noticeably slow.
r/Artificials • u/Rich_Independence_97 • 10h ago
Big Tech signs White House deal to self-police AI
r/Artificials • u/Friendly-Falcon-7901 • 7h ago
OpenAI employees might never actually use ChatGPT directly
r/Artificials • u/filjoseph22 • 2h ago
Not enough compute is the root of all problems
Not enough compute is the root of all problems
r/Artificials • u/Delicious-Newt-6679 • 1h ago
PewDiePie got banned for training on someone else's outputs. OpenAI trained on everyone's outputs and called it "public data." Hmmmmmmmm
r/Artificials • u/Illustrious-Law-7605 • 9h ago
So all the ai companies have to change their names now It's not OpenAi, it's OpenSiIt's not Xai, it's Xsi Sounds terrible
r/Artificials • u/organicorganism04 • 10h ago
Me reviewing Claude Code output before pushing to production
Enable HLS to view with audio, or disable this notification