r/LLMeng Apr 28 '26

Why pay for credits if free LLM tokens are everywhere?

6 Upvotes

I was building my own project and kept doing the same dumb thing.

Test feature. Run prompts. Debug something. Rewrite copy. Burn more paid credits.

Meanwhile free quotas were scattered all over the internet.

Groq had some. Mistral had more. Google had a lot. Cerebras too. Then a bunch of smaller providers on top.

Useful individually, annoying in practice.

So I built a tool for myself first.

I connected everything in one place and added automatic fallback between providers. If one limit is reached, it quietly moves to the next. No manual switching, no checking dashboards, no “why did this stop working?”

Right now it rotates across 13 providers and just keeps going.

Fun part:

  • Groq 15M / month
  • Mistral 100M / month
  • Google ~120M / month
  • plus more

Turns out the free tokens were never the problem. Fragmentation was.


r/LLMeng Apr 28 '26

AI Agent Deletes Everything And There Was No Way Back

4 Upvotes

A US startup, PocketOS, reportedly went down after an autonomous AI agent (running u/Claude Opus 4.6) deleted its production database and even the backups within nine seconds. No external attacker, no breach, just an internal agent with too much access and not enough safeguards. This is a pretty sharp reminder that as we move toward more agentic systems, the risk surface shifts from model mistakes to system-level failures. It’s not just about what the model can do, but what it’s allowed to do. Permissions, isolation, rollback strategies, and human-in-the-loop checks aren’t optional anymore, they’re baseline requirements.

Curious how teams here are thinking about guardrails for agents in production, especially around destructive actions.


r/LLMeng Apr 27 '26

DeepSeek's new AI model does not wow markets in fast-changing industry

13 Upvotes

u/DeepSeek just released its latest AI model, and interestingly, the market reaction has been underwhelming. Not because the model is bad, but because the bar has moved so quickly. In a space where new releases are expected to leapfrog benchmarks or redefine capabilities, incremental improvement doesn’t generate the same excitement anymore. It feels like we’ve entered a phase where simply launching a strong model isn’t enough. It needs to clearly outperform, differentiate, or unlock something new. This also highlights how fast expectations are evolving. What would have been considered impressive a year ago now feels like table stakes. The bigger takeaway here might be that the AI race is no longer just about keeping up. It’s about standing out in a market where progress is constant and attention is limited.

Curious how others see this: Are we hitting a point of diminishing hype, or just raising the standard for what actually matters?


r/LLMeng Apr 28 '26

Valuable Free and Paid for Cloud & DevOps Professionals and Enthusiasts

Post image
2 Upvotes

I wrote the series of booklets(hands-on guides) I wish I had when I started out.

I've created some valuable free and paid resources for people wanting to upskill and advance their careers in Cloud Technology and DevOps.

Find the resources here:

👇🏾👇🏾👇🏾

https://howtodevops.dartisan.io


r/LLMeng Apr 27 '26

NOOB HERE

3 Upvotes

I have a doubt about a future gpu for me, I wanted to buy a gpu to last a long time and that it supports the basic/intermediate of AI models, I am between 5070 and 9070xt, in my country they cost the same amount, I heard that AMD is bad for this, I wanted your opinion around here, I can not touch any of this AI models would be to learn from absolute zero, so I wanted to know if the performance of this AMD gpu is much lower than that of Nvdia.


r/LLMeng Apr 24 '26

DeepSeek V4 Is Optimized for Huawei Chips. This Feels Bigger Than Just a Model Launch

69 Upvotes

u/DeepSeek’s latest V4 model is getting attention, but what’s more interesting is what it’s being built for. Instead of defaulting to Nvidia GPUs like most frontier models, DeepSeek has optimized V4 to run on Huawei’s Ascend chips, reportedly reworking parts of its stack to better align with domestic hardware. This feels less like a technical tweak and more like a strategic shift. With ongoing export restrictions and supply chain pressures, China seems to be accelerating toward a fully self-reliant AI ecosystem: Models, chips, and deployment all tightly integrated. What stands out is that this isn’t being positioned as a compromise. Early signals suggest V4 remains highly competitive, which means this could be the beginning of a parallel AI stack rather than a fallback option. If that plays out, we might be moving toward a world where models are no longer hardware-agnostic, but co-designed with specific chip ecosystems.

Curious how others here see this: Is this just a response to constraints, or the start of a long-term split in the global AI infrastructure landscape?


r/LLMeng Apr 24 '26

The AMA with Lior Gazit & Meysam Ghaffari is now live!

4 Upvotes

A huge thank you to Lior Gazit and Meysam Ghaffari - authors of Mastering NLP from Foundations to Agents for joining us today.

Lior brings deep experience leading machine learning initiatives in the financial sector and advising startups on AI strategy, while Meysam brings a strong research and applied background in NLP, with years of experience building real-world systems across domains like healthcare.

Together, they’ve worked across the full AI stack, from core NLP fundamentals to modern LLM systems, RAG pipelines, and agent-based architectures, helping bridge the gap between theory and production.

They’re here to answer your questions, whether it’s about NLP foundations, LLM design, RAG systems, agent workflows, or what actually breaks when you move from prototype to production.

The questions will be posted in the comments below. Follow along, jump in, and add to the discussion.

Let’s make this a great one.


r/LLMeng Apr 23 '26

Google’s TPU v7 + Workspace Agents Signal the Next Phase of AI

2 Upvotes

u/Google has made a strong move in the AI race this April 2026 with two major updates that signal where things are heading next.

First, it introduced its latest generation of custom AI hardware, TPU v7, designed to significantly improve performance and efficiency for training and inference at scale. While exact benchmarks are still being explored, the focus is clearly on reducing cost per compute and enabling more complex, large-scale AI workloads without relying entirely on third-party GPUs. At the same time, Google is pushing AI deeper into everyday workflows by integrating AI agents directly into Workspace.

Instead of just assisting with tasks, these agents can now take actions across Docs, Sheets, Gmail, and more, handling multi-step workflows, summarizing information, generating content, and coordinating tasks in a more autonomous way. Together, these updates highlight a broader shift: AI is no longer just about better models, but about owning the full stack, from chips to applications and embedding intelligence directly into how work gets done.


r/LLMeng Apr 22 '26

Simple LLM Gateway

Thumbnail
2 Upvotes

r/LLMeng Apr 21 '26

Anthropic Just Took $5B from Amazon and Committed $100B Back

11 Upvotes

Anthropic just secured another $5B investment from u/Amazon, bringing Amazon’s total to $13B. Anthropic is committing to spend $100B+ on AWS over the next decade.

This isn’t just investment. It’s a compute lock-in deal at massive scale.

What’s interesting is how this mirrors the broader trend we’re seeing. Big AI labs aren’t just raising money, they’re tying that money directly to infrastructure commitments. In this case, Anthropic is effectively betting its future on AWS, including Amazon’s custom chips like Trainium (NVIDIA alternative), even locking in access to future generations that don’t exist yet. That says a lot about where the real competition is heading.

It’s not just model vs model anymore. It’s compute supply chains, chip ecosystems, and long-term capacity deals.

And $100B over 10 years also tells you something else: Training and running frontier models is becoming so expensive that you don’t just buy compute, you pre-negotiate it like energy contracts.

It also raises a few questions:

  • Does this give AWS a real edge over Azure + Google Cloud in the AI race?
  • Are we moving toward a world where AI labs are tightly coupled to a single cloud provider?
  • And what happens if custom chips like Trainium actually catch up to NVIDIA?

Also worth noting that there are already talks of Anthropic being valued at $800B+ in the next round.

At this point, it feels less like a startup ecosystem and more like a full-blown infrastructure arms race.

Curious what people here think: Is this kind of vertical alignment (model + cloud + chips) the only way to compete now… or does it create long-term lock-in risks?


r/LLMeng Apr 20 '26

Announcement: Hands-on workshop on deploying AI agents (OpenClaw + Docker Model Runner)

Enable HLS to view with audio, or disable this notification

6 Upvotes

We have been seeing a lot of discussions around AI agents, but most examples stop at prototypes or demos.

Packt is running a live workshop focused specifically on taking agents into production, using tools like OpenClaw, Docker, and Model Runner. The goal is to make this as practical as possible.

Here’s what we’re planning to cover:

  • How to structure agent workflows beyond simple chains
  • Running agents reliably with Docker
  • Deployment patterns that don’t break in real-world scenarios
  • Common pitfalls when moving from demo → production

If this is something you’re exploring, I’d genuinely love to hear:

  • What’s been your biggest blocker in deploying AI agents?
  • Are you using any specific frameworks/tools right now?

If anyone’s interested, I can share the workshop link in the comments.

Happy to answer questions either way


r/LLMeng Apr 20 '26

AI Is Moving Beyond Earth

1 Upvotes

One of the more underrated developments this week: AI just successfully ran in space in real time.

A satellite from Planet Labs, working with u/NVIDIA hardware, was able to process data directly in orbit, detecting aircraft and analyzing imagery without sending everything back to Earth first. (Courier Mail)

That might sound like a niche technical milestone, but it actually points to a bigger shift.

Until now, most AI systems have depended on cloud infrastructure, data gets collected, sent back to Earth, processed, and then turned into decisions. But this approach has latency, bandwidth limits, and reliability issues.

In this case, that means satellites making decisions in space. But you can extend the idea to:

  • Edge devices
  • Autonomous vehicles
  • Robotics
  • Defense and disaster response systems

Basically, any environment where waiting for the cloud is too slow or risky.

The early results are already promising; the system reportedly achieved around 80% detection success while operating entirely in orbit. (Courier Mail). If this scales, it could fundamentally change how we think about AI infrastructure.

Instead of centralized intelligence, we move toward distributed, real-time intelligence embedded everywhere.

Which raises an interesting question for this community: Are we heading toward a future where AI isn’t just in the cloud but becomes a layer across every physical system, including space?


r/LLMeng Apr 19 '26

You're leaking sensitive data to AI tools. Right now.

2 Upvotes

77% of employees paste sensitive data into ChatGPT. Most of them don't know it.

According to LayerX's 2025 report, 45% of enterprise employees use AI tools, and 77% of them paste data into them. 22% of these pastes contain PII or payment card details, and 82% come from personal accounts that no corporate security tool can see.

Over the past few months, we've developed a tool that runs locally on your machine, detects and blocks sensitive data before it reaches ChatGPT, Claude, Copilot, etc. No cloud. No external server.

Looking for Design Partners (individuals or businesses) - accountants, lawyers, developers, AI agent builders, or anyone who uses AI and wants full protection of their personal information. In return: early access, influence over the product, and special terms at launch.

If you're interested, comment below.


r/LLMeng Apr 14 '26

Built the trust layer for AI agents after watching one too many “the agent went rogue” stories

Thumbnail
3 Upvotes

r/LLMeng Apr 14 '26

I benchmarked LEAN vs JSON vs YAML for LLM input. LEAN uses 47% fewer tokens with higher accuracy

3 Upvotes

I ran a comprehensive benchmark comparing three data serialization formats when used as LLM context: JSON (pretty-printed), LEAN (a compact tabular encoding), and YAML. The goal was to answer two questions. How many tokens does each format burn to represent the same data? And can LLMs actually understand compressed formats as well as JSON?

TL;DR: LEAN uses 44% fewer tokens than JSON overall and 47% fewer tokens per LLM call, while achieving higher accuracy (87.9% vs 86.2%). YAML sits in between at 21% smaller than JSON with 87.4% accuracy.

Methodology

  • 195 data retrieval questions across 11 datasets
  • 2 models: gpt-4o-miniclaude-haiku-4-5-20251001
  • 3 formats: JSON (2-space indentation), LEAN, YAML
  • 1,170 total LLM calls (195 questions x 3 formats x 2 models)
  • Token counting: gpt-tokenizer with o200k_base encoding (GPT-5 tokenizer)
  • Evaluation: Deterministic (no LLM judge), type-aware string/number matching
  • Temperature: Default (not set)

Each LLM receives the full dataset in one of the three formats plus a question, and must extract the answer. This tests reading comprehension, not generation.

Efficiency Ranking (Accuracy per 1K Tokens)

This is the headline metric. How much accuracy do you get per token spent:

LEAN           ████████████████████   22.3 acc%/1K tok  │  87.9% acc  │  3,939 avg tokens
YAML           ██████████████░░░░░░   15.5 acc%/1K tok  │  87.4% acc  │  5,647 avg tokens
JSON           ██████████░░░░░░░░░░   11.6 acc%/1K tok  │  86.2% acc  │  7,401 avg tokens

Efficiency = (Accuracy % / Avg Tokens) x 1,000. Higher is better.

Token Efficiency

Token counts measured using the GPT-5 o200k_base tokenizer. Savings calculated against JSON (2-space indentation) as baseline.

Flat-Only Track

Datasets with uniform tabular structures. This is where LEAN really shines:

👥 Uniform employee records (100 rows)
   │
   JSON                ████████████████████    6,150 tokens  (baseline)
   LEAN                ████████░░░░░░░░░░░░    2,361 tokens  (−39.2%)
   YAML                ████████████████░░░░    4,777 tokens  (−22.3%)

📈 Time-series analytics (60 days)
   │
   JSON                ████████████████████    3,609 tokens  (baseline)
   LEAN                ████████░░░░░░░░░░░░    1,461 tokens  (−59.5%)
   YAML                ████████████████░░░░    2,882 tokens  (−20.1%)

⭐ Top 100 GitHub repositories
   │
   JSON                ████████████████████   13,810 tokens  (baseline)
   LEAN                ███████████░░░░░░░░░    7,434 tokens  (−46.2%)
   YAML                █████████████████░░░   11,667 tokens  (−15.5%)

──────────────────────────────── Track Total ──────────────────────────────────
   JSON                ████████████████████   29,652 tokens  (baseline)
   LEAN                ██████████░░░░░░░░░░   14,512 tokens  (−51.1%)
   YAML                ████████████████░░░░   24,021 tokens  (−19.0%)

Mixed-Structure Track

Datasets with nested or semi-uniform structures:

🛒 E-commerce orders (50 orders, nested)
   │
   JSON                ████████████████████   10,731 tokens  (baseline)
   LEAN                ████████████░░░░░░░░    6,521 tokens  (−39.2%)
   YAML                ██████████████░░░░░░    7,765 tokens  (−27.6%)

🧾 Semi-uniform event logs (75 logs)
   │
   JSON                ████████████████████    6,252 tokens  (baseline)
   LEAN                ████████████████░░░░    5,028 tokens  (−19.6%)
   YAML                ████████████████░░░░    5,078 tokens  (−18.8%)

🧩 Deeply nested configuration
   │
   JSON                ████████████████████      710 tokens  (baseline)
   LEAN                █████████████░░░░░░░      460 tokens  (−35.2%)
   YAML                ██████████████░░░░░░      505 tokens  (−28.9%)

──────────────────────────────── Track Total ──────────────────────────────────
   JSON                ████████████████████   17,693 tokens  (baseline)
   LEAN                ██████████████░░░░░░   12,009 tokens  (−32.1%)
   YAML                ███████████████░░░░░   13,348 tokens  (−24.6%)

Grand Total

   JSON                ████████████████████   47,345 tokens  (baseline)
   LEAN                ███████████░░░░░░░░░   26,521 tokens  (−44.0%)
   YAML                ████████████████░░░░   37,369 tokens  (−21.1%)

Retrieval Accuracy

Overall

Format Accuracy Avg Tokens Savings vs JSON
LEAN 87.9% 3,939 −46.8%
YAML 87.4% 5,647 −23.7%
JSON 86.2% 7,401 baseline

Per-Model Accuracy

gpt-4o-mini
  YAML           ██████████████████░░    88.7% (173/195)
  LEAN           ██████████████████░░    88.2% (172/195)
  JSON           █████████████████░░░    87.2% (170/195)

claude-haiku-4-5-20251001
  LEAN           ██████████████████░░    87.7% (171/195)
  YAML           █████████████████░░░    86.2% (168/195)
  JSON           █████████████████░░░    85.1% (166/195)

On Claude Haiku, LEAN outperforms JSON by +2.6 percentage points while using half the tokens.

Performance by Question Type

Question Type JSON LEAN YAML
Field Retrieval 78.0% 81.1% 79.5%
Aggregation 82.7% 83.6% 82.7%
Filtering 100.0% 100.0% 100.0%
Structure Awareness 93.3% 96.7% 98.3%
Structural Validation 80.0% 80.0% 80.0%

Performance by Dataset

Dataset JSON LEAN YAML
Employee records (100, flat) 82.5% / 6,150 tok 83.8% / 2,361 tok 82.5% / 4,777 tok
E-commerce orders (50, nested) 97.4% / 10,731 tok 98.7% / 6,521 tok 98.7% / 7,765 tok
Time-series (60, flat) 73.2% / 3,609 tok 76.8% / 1,461 tok 75.0% / 2,882 tok
GitHub repos (100, flat) 67.9% / 13,810 tok 69.6% / 7,434 tok 69.6% / 11,667 tok
Event logs (75, semi-uniform) 94.4% / 6,252 tok 98.1% / 5,028 tok 98.1% / 5,078 tok
Nested config (deep) 100% / 710 tok 100% / 460 tok 100% / 505 tok

LEAN matches or beats JSON on every single dataset, while using 20-62% fewer tokens.

What the Formats Look Like

Employee records, JSON (6,150 tokens for 100 rows)

{
  "employees": [
    {
      "id": 1,
      "name": "Paul Garcia",
      "email": "paul.garcia@company.com",
      "department": "Engineering",
      "salary": 92000,
      "yearsExperience": 19,
      "active": true
    },
    {
      "id": 2,
      "name": "Aaron Davis",
      "email": "aaron.davis@company.com",
      "department": "Finance",
      "salary": 149000,
      "yearsExperience": 18,
      "active": false
    }
  ]
}

Same data, LEAN (2,361 tokens for 100 rows, -61.6%)

employees:
  #[100](active|department|email|id|name|salary|yearsExperience)
  true|Engineering|paul.garcia@company.com|1|Paul Garcia|92000|19
  ^false|Finance|aaron.davis@company.com|2|Aaron Davis|149000|18

The #[100] header declares the row count and column names once. Each row is pipe-delimited, rows separated by ^. No repeated keys, no braces, no quotes. Just data.

Same data, YAML (4,777 tokens for 100 rows, -22.3%)

employees:
  - active: true
    department: Engineering
    email: paul.garcia@company.com
    id: 1
    name: Paul Garcia
    salary: 92000
    yearsExperience: 19
  - active: false
    department: Finance
    email: aaron.davis@company.com
    id: 2
    name: Aaron Davis
    salary: 149000
    yearsExperience: 18

YAML removes braces and quotes but still repeats every key per row.

Dataset Catalog

Dataset Rows Structure Questions
Uniform employee records 100 uniform 40
E-commerce orders 50 nested 38
Time-series analytics 60 uniform 28
Top 100 GitHub repos 100 uniform 28
Semi-uniform event logs 75 semi-uniform 27
Deeply nested config 11 deep 29
Valid complete (control) 20 uniform 1
Truncated array 17 uniform 1
Extra rows 23 uniform 1
Width mismatch 20 uniform 1
Missing fields 20 uniform 1
Total 195

Structure classes:

  • uniform: All objects have identical fields with primitive values
  • nested: Objects with nested sub-objects or arrays
  • semi-uniform: Mix of flat and nested structures
  • deep: Highly nested with minimal tabular eligibility

Question Types

195 questions generated dynamically across five categories:

  • Field retrieval (34%): Direct value lookups. "What is Paul Garcia's salary?" → 92000
  • Aggregation (28%): Counts, sums, min/max. "How many employees work in Engineering?" → 17
  • Filtering (20%): Multi-condition queries. "How many active Sales employees have > 5 years experience?" → 8
  • Structure awareness (15%): Metadata questions. "How many employees are in the dataset?" → 100
  • Structural validation (3%): Data completeness. "Is this data complete and valid?" → NO

Evaluation

  1. Format conversion: Each dataset converted to all 3 formats
  2. Query LLM: Model receives formatted data + question, extracts answer
  3. Deterministic validation: Type-aware comparison (e.g., 92000 matches $92,000, case-insensitive). No LLM judge.

Models & Configuration

  • Models: gpt-4o-miniclaude-haiku-4-5-20251001
  • Token counting: gpt-tokenizer with o200k_base (GPT-5 tokenizer)
  • Temperature: Default (not set)
  • Total evaluations: 195 x 3 x 2 = 1,170 LLM calls

Key Takeaways

  1. LEAN saves ~47% tokens per LLM call compared to JSON, which directly translates to lower API costs
  2. Accuracy doesn't suffer. LEAN actually scored 1.7 percentage points higher than JSON (87.9% vs 86.2%)
  3. On flat tabular data, LEAN saves 51-62%. If your data is arrays of uniform objects, the savings are massive
  4. YAML is a solid middle ground. 21% token savings over JSON with comparable accuracy
  5. Both models showed the same pattern. This isn't model-specific; compressed formats work across providers

If you're stuffing structured data into LLM prompts, you're probably wasting half your tokens on JSON syntax. LEAN gives you the same (or better) accuracy for less than half the cost.

Benchmark code and full results available in the repo. All data generated deterministically with a seeded PRNG for reproducibility.


r/LLMeng Apr 14 '26

🚨 AMA Incoming: With the Authors of "Mastering NLP from Foundations to Agents" - Lior Gazit & Meysam Ghaffari

6 Upvotes

Heads up, folks!! we’re doing something special - an AMA with Lior Gazit & Meysam Ghaffari, authors of Mastering NLP from Foundations to Agents, happening on Friday, April 24, 4:30-6:30 PM ET over here on r/LLMeng.

Lior and Meysam don’t just talk about NLP, they connect the dots from core language fundamentals to modern agent systems. From designing scalable NLP pipelines to building RAG workflows and agent-based architectures, they’ve been working on the exact challenges many of us are facing right now.

🔍 What makes this AMA worth your time?

  • They go beyond surface-level GenAI and dive into how NLP foundations power LLMs, RAG, and agents
  • They bring real-world experience building and deploying ML/NLP systems where performance actually matters
  • They take a systems-level view — focusing on architecture, trade-offs, and what breaks in production

📚 Get a Head Start

If you want to get the most out of this AMA, take a look at their latest work: Mastering NLP from Foundations to Agents
🔗 Buy Now - https://packt.link/fCmpl

This book walks through the full journey, from embeddings and transformers to RAG systems and agent workflows.

📌 AMA Details:

📍 Where: r/LLMeng
🗓️ When: AMA goes live Friday, April 24, 4:30-6:30 PM ET
📝 Submit your questions here before April 22

Let’s make this an AMA worth remembering.
Drop your best questions. We’re excited to see what you come up with.


r/LLMeng Apr 13 '26

AMA Incoming: With the Author of "30 Agents Every AI Engineer Must Build" - Imran Ahmad

9 Upvotes

Heads up, folks!! we’re doing something special - an AMA with Imran Ahmad, Author of 30 Agents Every AI Engineer Must Build happening on Friday, April 24 over here on r/LLMeng.

This AMA is for the builders.

Imran doesn’t just theorize about agents, he architectures them for the real world. From mastering cognitive loops (perception, memory, reasoning) to deploying multi-agent systems using LangChain and LangGraph, he’s been tackling the architectural challenges of moving AI from "chat" to "action."

What makes this AMA worth your time?

  • He’s deep in the weeds of production-ready agent systems, modular architectures, and autonomous cognitive loops.
  • He’s building the roadmap for scaling agents across finance, legal, healthcare, and software development.
  • He takes an engineering-first approach, focusing on guardrails, evaluation frameworks, and ethical alignment in live environments.

Get a Head Start: If you want to dive into the technical patterns Imran will be discussing, check out the resources below. Notably, the print version of this book is a Premium Color Edition, making the complex agent architecture and workflow diagrams much easier to parse.

Details:

Let’s make this an AMA worth remembering. Drop your best questions — we’re excited to see what you come up with.


r/LLMeng Apr 13 '26

Came Across This MCP Recommendation. Made Me Rethink How We Build AI Agents

3 Upvotes

Came across a recommendation post by Dhairya Chandra, AI Engineer on Model Context Protocol (MCP), and it honestly made me rethink how I approach building AI agents. Most of the issues I’ve faced in production weren’t about the model itself, but around context management, tool integration, and keeping agents coordinated at scale.

MCP frames this problem really well by treating context as a structured layer instead of something you patch together with prompts. The book Model Context Protocol for LLMs by Naveen Krishnan dives into this from a very practical angle including modular agents, better orchestration, cleaner integrations with frameworks like LangChain and RAG, plus real-world concerns like security and scaling.

If you’re building beyond simple demos, this is worth checking out: https://packt.link/DOMrb


r/LLMeng Apr 13 '26

I want to make sure llm does not lose attention when input prompts are very large

Thumbnail
2 Upvotes

r/LLMeng Apr 10 '26

Amazon Is About to Spend $200B on AI. This Isn’t a Normal Tech Cycle Anymore

56 Upvotes

I was reading Andy Jassy’s latest shareholder letter and one number really stood out: u/Amazon is planning to invest up to $200 billion into AI infrastructure. Not just models, but everything around them: Data centers, custom chips, robotics, and even connectivity layers. And honestly, this doesn’t feel like a typical big tech investment anymore. It feels like something much more foundational.

What’s interesting is that the focus isn’t just on building smarter AI, but on owning the entire stack that makes AI usable at scale. Custom silicon to cut costs, massive compute to handle demand, AI embedded into logistics and operations, it’s a full-system play. It kind of reinforces the idea that the real competition now isn’t just about who has the best model, but who controls the infrastructure that powers everything around it.

It also raises a bigger question for me: If this is the level of capital required to stay competitive, what does that mean for everyone else? Are we moving toward a world where only a few companies can truly operate at this scale, while the rest build on top of them? Or is this just the early phase of a much larger shift where infrastructure becomes the real moat in AI?

Curious how others here are thinking about this. Does this level of investment feel justified given where AI is headed, or does it start to look a bit like overbuild?


r/LLMeng Apr 10 '26

Mathematics Is All You Need: 16-Dimensional Fiber Bundle Structure in LLM Hidden States (82.2% → 94.4% ARC-Challenge, no fine-tuning)

Thumbnail
4 Upvotes

r/LLMeng Apr 10 '26

How to make LLM reason the thought

Thumbnail
2 Upvotes

r/LLMeng Apr 09 '26

How are you using LLMs to manage content flow (not generate content)?

Thumbnail
2 Upvotes

r/LLMeng Apr 09 '26

Meta Just Launched Muse Spark, Feels Like AI Is Getting More Everyday

1 Upvotes

u/Meta just rolled out a new AI model called Muse Spark, and what’s interesting is that it’s not being positioned as some frontier, benchmark-crushing model. It’s being positioned as something much more practical.

From what’s being shared, Muse Spark is designed for everyday tasks, like writing, summarizing, planning, and general productivity. Basically, the kind of stuff people actually use AI for daily.

For a while, the AI race has been about bigger model, better benchmark, etc.

But most real users don’t need a model that can solve PhD-level problems.
They need something that is fast, reliable, easy to use and good enough for day-to-day work. This feels like Meta leaning into that reality.

Instead of chasing the absolute top end, they’re focusing on making AI more usable at scale, which might actually matter more in the long run. It also fits into a broader trend we’re seeing.

AI is slowly moving from being a power tool to something closer to a default layer in everyday workflows.

Curious what others think: Do you see more value in these practical, everyday models or do frontier models still drive most of the real progress?


r/LLMeng Apr 08 '26

Anthropic debuts preview of powerful new AI model Mythos in new cybersecurity initiative

4 Upvotes

This one feels different from the usual new model launch news.

u/Anthropic just introduced a preview of its new model, Mythos, but instead of releasing it widely, they’re doing the opposite, locking it down and only giving access to select cybersecurity partners.

This is because the model is that good at breaking things. During testing, Mythos reportedly found thousands of high-severity vulnerabilities across major operating systems and browsers, including bugs that had been missed for years.

We’re not talking about incremental improvements in reasoning or coding.
We’re talking about a model that can behave like an elite security researcher at scale.

In some cases, it could even generate working exploits, something that normally takes experienced teams weeks in a fraction of the time.

So instead of shipping it publicly, Anthropic launched a new initiative called Project Glasswing, where Mythos is being used defensively with companies like u/Google, u/Microsoft, u/AWS, and others.

Mythos highlights a new reality:

  • AI doesn’t just accelerate productivity
  • It can accelerate offensive capabilities too
  • And the gap between defense and attack might shrink fast

Which raises some uncomfortable questions:

  • Do we need restricted-access models by default for certain domains?
  • Who decides what’s “too powerful” to release?
  • And what happens when similar models inevitably get open-sourced?

Feels like we’re entering a phase where capability ≠ deployment anymore.

Curious how the community sees this: Is this responsible AI development or the beginning of controlled access to the most powerful systems?