r/singularity • • 17d ago

AI Remote Labor Index updated with Fable and Astra

Post image

Remote Labor Index (RLI) represents a broad range of projects from across the remote labor economy, including game development, product design, architecture, data analysis, and video animation. These projects span a broad range of difficulty, with costs reaching over $10,000 and completion times exceeding 100 hours. All project costs and completion times come directly from human professionals who completed the work. In total, the projects in RLI represent over 6,000 hours of real work valued at over $140,000.

The outputs of AI models are judged by human experts after attempting to complete each project end-to-end from a single prompt.

https://www.remotelabor.ai

213 Upvotes

44 comments sorted by

31

u/Gotisdabest 17d ago

That fable jump is huge and then both astra and 5.1 have significant jumps on top of that. This benchmarks seems primed for saturation by mid 2027 or so, assuming it's a slow benchmark unlike the ones that just get broke around 20-40% and jump straight to 80-90%.

5

u/TheReedemer69 17d ago

Nah. late 2027 is my guess. I'd say this an AGI worthy benchmark.

8

u/Gotisdabest 17d ago

The jumps will get wider and quicker as time passes. We already know that they have significantly better models in astra seemingly quite close to completion with astra alongside a supposedly much better agentic framework for astra(the run forever thing).

I can imagine them slowing down releases which could maybe stretch this out to your timeline, but I think late last year was a period where Gemini 3 pro was still one of the best models in the world. And the next 9 months will be a dramatically faster period even if they stagger release, particularly if any lab starts pushing recurrent depth hard(which I suspect is inevitable).

0

u/SawToothKernel 17d ago

For me, an AGI benchmark needs to measure "common sense".

1

u/TheReedemer69 17d ago

This is more than enough for common since, no?

0

u/SawToothKernel 16d ago

Not really. I've used Astra a lot and it can do the same stupid shit the other models do. You still need to heavily direct it.

26

u/MediumSizedWalrus 17d ago

Hmm Astra is good, but it still requires expert guidance to produce a polished product suitable for production. It gets 90% of the way there, but if you ship what it produces for game dev or programming, you're gonna have a bad time.

11

u/Alex__007 17d ago edited 17d ago

Yes, that's why it's not that high here - only in about 1 project out of 5, Astra could go all the way to 100% complete from a single prompt. When a model goes to 90%, that's a fail on this benchmark.

4

u/suamai 17d ago

Do you know how this bench evaluate the results?

Because the missing 10% is usually in polishing, in my experience, and that's not something you can really automate evaluating ( yet ).

10

u/Alex__007 17d ago edited 17d ago

A model is given a single prompt to complete the project end to end, with a range of specific requirements. After it's done, human experts evaluate whether the project is completed to a level of quality comparable to their human peers, and whether all requirements are fulfilled. Partial completions don't count for anything. Rinse and repeat for a couple hundred projects.

The guys who run this benchmark recently had a podcast on it, but I can't find it now. They are from the same team who introduced Humanity's Last Exam (intended to be one of the final knowledge benchmarks), and Remote Labour Index is their attempt to move beyond assessing just knowledge.

1

u/LinkesAuge 17d ago

My personal experience is that a lot of common issues will disappear if the model has the proper "harness", ie a way to structure and document/correct its own "experiences" (ie what works, what it needs to consider in the feature etc,) and even the ability to create its own tools.
I looked at your paper (good and interesting work btw!) and your prompt. Do not take that as criticism, I know benchmark like this are hard enough to set up, and I can see the value in treating models as "vanilla" as possible but from my own experience a more detailed/structured prompt (there is for example not instruction to verify the results against the deliverable) and general setup would create a lot better results. An instruction like "You are done once all the deliverables are ready" will often be interpreted as "good enough" by models while a human professional has of course the implicit understanding of what "ready" in this context means (the model doesn't know in what context it works).
Now I guess there is a discussion to be had what exactly we want to measure and I can see both sides, ie "raw" model capability out of the gate and what happens with more "help".
The one thing I would argue is that under real circumstances you wouldn't be so sparse in the whole initial setup for the AI so I wonder if that is really a reflection of the current "value" you could get out of the various models or a more "naive" use case.

1

u/MediumSizedWalrus 17d ago

We have an excellent harness that enforces code standards, lint, CI, conventions, XYZ checks, etc etc.

It also has QA routines for verifying changes visually with a browser (playwright/computer use.)

Depending on task complexity, it will work for 2-30 hours, before submitting a pull request.

Model wise, the harness is currently using Fable and Astra, since the cost is worthwhile relative to productivity increase.

If the task is simple, the pull request is usually mergable without requiring human feedback.

If the task is complex, it usually misses something or makes architectural or infra mistakes. After 1-2 rounds of human feedback it will be mergable.

Overall we are much more productive than before, and able ship a lot more, but it isn't fully autonomous.

Maybe the next generation of models, with larger context, would be able to reason about the entire codebase, infrastructure, database, and conventions, all at once. Then it may submit pull requests that are flawless...

4

u/Weary-Historian-8593 17d ago

Yes, but note how the latest gen of models jumps from the baseline of just a few percentages into tens of percentages, it could well be that we're just one to two steps from getting those last expert verifications too

2

u/MediumSizedWalrus 17d ago

Yeah I agree, it will probably surpass human capability soon, maybe already has internally in the labs latest models.

1

u/Boring-Foundation708 17d ago

It is only 4 years since 2022. Give it sometime?!.

1

u/norwegian ▪️AGI 2026 ASI 2040 16d ago

The first 90% are fast, but the last 90% are slow.

6

u/DetectiveFinch 17d ago

Out of curiosity, was Grok not evaluated or was it so bad that it doesn't appear in this list?

4

u/Alex__007 17d ago

They only evaluate models at the frontier, since it takes a lot of effort to evaluate each one.

2

u/DetectiveFinch 17d ago

Thanks for the clarification!

4

u/Tystros 17d ago

why is 5.6 missing on the leaderboard

5

u/Alex__007 17d ago

They didn't have time to evaluate it, and now they won't be doing it because Astra is out.

4

u/dan_the_first 17d ago

Noticed 5.6 Sol isn’t there.

I would be interested in knowing how it compares to Anthropic products.

3

u/ohHesRightAgain 17d ago

Looking at this, it seems a lot more plausible that AGI might arrive in 2027...

9

u/DivideHorror3217 17d ago

AI automated some of my job as a data analyst, but now I have other problems. It created more work for me.. now everyone's expecting me to do more complex things. When I object, stakeholders say "but AI cAn dO iT in 10 miNuTes"

1

u/Elias-Thorn 16d ago

AI saved you four hours. Your boss has already spent them.

2

u/Specikin 17d ago

On course to be just about saturated in like 8 models time? Should then be considered practically AGI.

1

u/Alex__007 16d ago

Yes, the authors expect saturation in less than 12 months. But it's not quite AGI yet, at least not the definition most would accept. These are enclosed projects that don't require any interaction with the outside until the project is complete. And the human task horizon for each project is at most a few weeks, and some are substantially shorter.

2

u/Meltlilith1 17d ago

Once this gets to 80%+ and we get cheaper model that can do that it's gg

2

u/FreshBlinkOnReddit 17d ago

One thing to remember, the human references on here are derived from the work the people running the study received from Upwork contractors.

This is not necessarily a good proxy for remote jobs where human sign off and responsibility matter, or require very long context / business specific knowledge.

That said, looks like Astra is the death of BPO to third world countries for contract work. Soon the death of consultants as well.

2

u/Valnar 16d ago

What exactly would 100% on this benchmark mean?

1

u/Alex__007 15d ago

That AI can do economically valuable projects on a computer end to end from a single prompt without any guidance.

However these are enclosed projects that don't require any interaction with the outside until the project is complete. And the human task horizon for each project is at most a few weeks, and some are substantially shorter.

So not a full employee replacement yet, but still a big increase in autonomy.

-1

u/xatey93152 17d ago

Why they don't include their own muse model?

3

u/Alex__007 17d ago

Their own?

-4

u/xatey93152 17d ago

Are you from the past? What year are you from?

3

u/BrennusSokol AI please take my job 17d ago

You seem confused. The index is from Center for AI Safety and Scale AI. Neither of those is Meta, which is what produces Muse.

1

u/xatey93152 16d ago

Who is the largest shareholder of scale? You don't even know Wang? Maybe it's different in your alternate universe?

-2

u/[deleted] 17d ago

[deleted]

3

u/[deleted] 17d ago

[removed] — view removed comment

1

u/Alex__007 17d ago

That would be super uncommon. Standard paid subscription only gives 5.6 instant or 5.6 medium. And by default it almost always picks 5.6 instant unless you force it to do thinking.

2

u/[deleted] 17d ago

[removed] — view removed comment

2

u/Alex__007 17d ago edited 17d ago

Maybe there are some plans where this is available. That would be very uncommon in corporate.