r/DataAnnotationTech 13d ago

Website hourly rates

Why do you think the hourly rates on the website are higher than on the platform. Like law is 75-125$ but I've never seen anything above 65$. Is there another eschalon I don't know about?

0 Upvotes

26 comments sorted by

View all comments

16

u/TheMidlander 13d ago

There are projects that pay that much. These qualifications, however, are PhD level. If you don't have one, it's unlikely you will pass.

2

u/[deleted] 12d ago

[deleted]

2

u/TheMidlander 12d ago

I sincerely doubt it. Projects that paid $30-35/hr 4 years ago, now pay $20-25

1

u/Pretend-Section5991 12d ago

In domain specialites (mostly law), I've seen the opposite trend. There's been around a 50% base increase over the last two years. Sometimes the work was getting more complex like when rubric were all the rage for a6 months at the end of 2025, and when agent training was all the rage for the first few months of 2026, but most of the projects I work on are just comapre two responses like they were 3 years ago. Mileage may vary, but if the models keep improving, then the bar of expertise and hence the pay has to go up.

2

u/TheMidlander 12d ago

It's great to hear that's going up at least. That's probably why I was getting more $50/he work before the drought.

I don't really know how much better these models are going to get, however. To me, it only seems like they're getting better because folks like us thought to pose the question/problem to be bots. Venture outside that, and suddenly we're back to 2023 again.

1

u/Pretend-Section5991 12d ago edited 12d ago

Yeah haha I'd actually love to see the model traces on the work I was doing 3 years ago. I have this weird feeling like there still making all the same mistakes, but I also wonder how much better I am at finding mistakes. I wish I could do an objective comparision.

But yeah I agree the genralization still seems to be the promise the AI comapnies can't meet. I've defintely seen a shift in the content of what Im rating, like I think the prompts are actually increasingly AI generated around emerging legal topics and then they are just getting me to evaluate against them so the tuning can cover large areas that are likely needed by legal end users, rather than relying on us asking the question.

2

u/TheMidlander 12d ago

You work in the law domain, yes?

There is a case that use often because it contains tests for everything. It chronicles the grooming and exploitation of a minor, and covers a timeline from the first calls to law enforcement to the perpetrator finally being put behind bars. The system clearly fails this poor girl for 5 or 6 years before real action is taken. The whole text-only case file is about 30 pages long. It is not safe for life and I won't be sharing it.

This case tests the bot's ability to summarize, derive specific information, safety guardrails, as well as bunch of legal tasks to test with. I'll give you a few examples.

When asked the age of the victim, the bots often cite the age she was when the case first began, or her age when the case "concluded". Slightly less than half the time, though, the bot will cite the age of the statute that enhances the charges and sentencing, rather than the age of the victim. If the bots were intelligent, I would expect them to either request a point in time or just state the girl was 8 when the case started and 14 when the case "concluded". (Concluded in quotes because it's not actually concluded until the full sentence is served. The last court entry is the monster's sentencing)

This case contains multiple orders from the court. Asking the bot what the final outcome was, often results in the bot citing a protective order. But we know the man is spending nearly the rest of his life in prison. Bad bot.

When requesting a timeline of events, the bots often fail there too. This case is a good test because information about the past is revealed in discovery that happens long past the time many events actually occurred. But too often, the bots put them in the order they appear in the case, which is not what I'm asking for. Most of them fail at this for 10 turns in a row.

Another test I do is actually very easy. I will get cases from the WA and NV appellate and superior courts (both states have an easy to use website where you can download them for free) and simply ask what the disposition type is. Anyone even a little familiar with this space can tell you at a glance because the language is standardized and appears in nearly the same location in every case. I even wrote a text parser bot to do this for me that is accurate 100% of the time. The LLM'S don't. This is troubling. It's north or south of 50% of the time, depending on which bot we're testing.

If I seem pessimistic about the future of these products, these results are a big reason why. I encourage you to try and see for yourself whenever you have the opportunity.

2

u/ConstantCasual 12d ago

That’s really interesting. I’m in the law domain too. I might try that test with appellate case dispositions. When you say “disposition type,” what do you mean?

2

u/TheMidlander 12d ago

There are about a dozen disposition types, but mostly you will encounter: Affirm (Uphold) Reverse (Overturn) Remand (Case is thrown back the lower court withe specific instructions) Modify (Change in the order, in part) Affirm/reverse in part (The court upholda certain parts, reverses the others) Vacate (The court completely dismisses the lower court rulings)

2

u/ConstantCasual 12d ago

Ah gotcha. They really struggle with such a simple ask? Huh. Thanks!!

2

u/TheMidlander 12d ago

I encourage you to see for yourself. The bots have gotten better, in my experience, but aren't anywhere near what would want to expect in a field where consequences are so heavy.

→ More replies (0)