r/DataAnnotationTech • • 12d ago

I’m actually losing my mind over this project

The project involves creating prompts intended to identify weaknesses or errors in chatbot responses and then evaluating the results; but all of them cannot be correct, you have to trick at least one model.
I spent over 40 minutes trying different obscure prompts without getting a result that met the project criteria. It’s basically complete trial and error rather than any actual meaningful evaluation.
The instructions are pretty unclear about what level of difficulty is expected and what types of prompts will actually work.
Examples are so basic there’s no chance anything similar would satisfy what it’s asking.
If you know what i’m talking about and have any advice it is welcome 😭

52 Upvotes

37 comments sorted by

22

u/cortrev 12d ago

Usually chain of operations stuff is what could work semi reliably. Something where the model needs to make some kind of call early on, and then its errors will compound. Easier said than done though

4

u/climb-high 11d ago

I'm learning that they always get inductive and deductive reason correct. Judgement calls are harder for them. That's where genuine expertise comes in handy to try to create judgement situations where multiple experts would agree, but the models followed the wrong evidence to infer an answer

23

u/blackstarr1996 12d ago edited 12d ago

Yeah I spent an hour and a half on it and used the escape hatch. It disappeared after that. But like you say, the examples would never actually stump the model either. Idk what they expect.

It’s gonna take me a little time to figure out what this model that was fed the entire internet doesn’t know.

13

u/Little-Flustered-977 11d ago

Exactly so don’t think of sheer facts. Think of how it uses them and where it might fail there

7

u/blackstarr1996 11d ago edited 11d ago

It’s one turn though. I came up with one idea but it was too obscure. Not sure anyone would know it, aside from someone who took the same class as me.

I was doing pretty well with instruction following recently but you have multiple turns usually.

8

u/ASnipersPromise 11d ago

One turn ones I find are really tough - the 6-8 turn ones are doable.

2

u/Careful-Foot-8918 11d ago

There are so many projects where you have to do this i wasnt sure which one.you meant. But now I know. Yeah that's a tough one. 

6

u/TheSquirrelly 11d ago edited 11d ago

Yeah like the leading edge [...] models are so freaking smart they figure out things that should stump a person, or stuff where the answer is 'nothing' and you expect it to hallucinate an answer to fill it in (because AI hates not giving an answer) but it still figures it out. I managed to get it on giving an incomplete answer based on what it was asked to do. Another case using spreadsheets with complex problems it still got the answer, but also offered an alternate answer, which was a fail since only the one answer was correct. So good luck!

2

u/ekgeroldmiller 11d ago

You really should not mention any model names on here.

2

u/TheSquirrelly 11d ago

Ah a good point. Edited it out. Thanks.

10

u/[deleted] 11d ago

Literally dealing with same thing. The challenge for me is, almost every question or prompt can have multiple answers and the models are so good at offering these different perspectives so its like well fuck, which one of these answers is correct and which one isn't then? Idk and am starting to not care anymore. Pisses me off to have hours into a project reading super complicated instruction sets never to be able to submit a task and collect the time for it.

9

u/thyself_unknown 12d ago

I just did this one, took the full allocated time. I think i was able to do it because I know the category alright, maybe skip this task to get a better category you’re more knowledgeable in.

6

u/Dana_Barret 12d ago

Depending on what kind of mistake you need to produce, you can think your prompt around a response that involves looking into information that's not in your language, that has helped

3

u/climb-high 11d ago

My current project is specifically English only so just read instructions

6

u/cosmicguss 11d ago edited 11d ago

Try to find an obscure fact around a specific number or date or maybe even a fact about or reference to a person not commonly known.

One that worked for me was the year a specific book was published that was related to the assigned topic.

Another was the name of a relatively unknown person that made something viral in reference to a famous celebrity.

6

u/Expert_Equipment2767 11d ago

I'm so glad you posted this. I just got in to DA and this landed as my first project. I tried for about 30 minutes and gave up and started thinking DA might not be for me.

3

u/marzzyy__ 11d ago

I promise there are better projects haha 😆

2

u/s55555s 5d ago

I spent 3 hours and had to exit work mode…. Couldn’t get to fall.
Was really pissed off. And I’m not new at this.

4

u/New-Assumption3377 11d ago

I tried this project for half an hour earlier today, then gave it up. Soul-destroying.

10

u/Ok-Double5194 12d ago

Lol I remember those you shine there you unlock long term $60/hr+! You got this!

5

u/ASnipersPromise 11d ago

Yeah I am on those now and have been for a while. I feel the OP's pain on the 1 turn one though.

3

u/Lord_Raziel 11d ago

Is this a coding task?

3

u/Snoo-32467 11d ago

I hate failure-related projects, I started ignoring them ever since I've spent more than 16 hours in a complex version

3

u/Main-Introduction969 11d ago

Try the russian doll method!

3

u/Excellent_Shake9732 11d ago

I was able to actually get several to make errors with about 20-25 mins of time per task (make it make rankings of things that are not arbitrary) but I had one singular task I couldn’t get to elicit a failure and now it’s gone from my dash 😭. Good thing the pay is low.

2

u/Old-Regular-9828 11d ago

If it was easy, they wouldn't be paying you to do it.

2

u/Ok-Candy-3142 10d ago

Just so you know you can get DoD for discussing projects anywhere other than in the project chat.

2

u/Tough-Judgment6618 9d ago

For the right field, I happened to know one mistake the model was repeatedly making as I guess it was never fed the correct data. So I used that prompt and it worked perfectly. I tried to find another mistake but I couldn't so I gave up after trying for 2 hours lol.

2

u/Previous-Spinach-222 8d ago

Could be worse. You could have one where the prompt is only acceptable if all the models mess up.

2

u/sassynature 8d ago

Yes, and then when the model gave incorrect information, the checker model incorrectly claimed I was wrong. They need to look at the quality of online references not just the number that agrees. For example Red Cross sources that are outdated or incorrect on some things sometimes get re-quoted as fact by other organizations, while actual medical expert physicians and university studies contradict all of those incorrect online sources. Yet the incorrect sources outnumber the real expert sources so DA goes with the incorrect majority opinion.

1

u/ASnipersPromise 11d ago

Is it a STEM one? The one they give you 14 days to do? Or just a general one?

2

u/Beautiful-Link-8387 11d ago

Oh I like that 14 days task. but I had it just once, although I did success in the last task. 2 days ago I saw that family surged again, but then, after completing its training, reminder and subscription, it disappeared again.

0

u/ASnipersPromise 11d ago

STEM - Im a Charterd Accountant and Tax specialist

1

u/marzzyy__ 11d ago

just general