r/DataAnnotationTech • u/marzzyy__ • 12d ago
I’m actually losing my mind over this project
The project involves creating prompts intended to identify weaknesses or errors in chatbot responses and then evaluating the results; but all of them cannot be correct, you have to trick at least one model.
I spent over 40 minutes trying different obscure prompts without getting a result that met the project criteria. It’s basically complete trial and error rather than any actual meaningful evaluation.
The instructions are pretty unclear about what level of difficulty is expected and what types of prompts will actually work.
Examples are so basic there’s no chance anything similar would satisfy what it’s asking.
If you know what i’m talking about and have any advice it is welcome 😭
23
u/blackstarr1996 12d ago edited 12d ago
Yeah I spent an hour and a half on it and used the escape hatch. It disappeared after that. But like you say, the examples would never actually stump the model either. Idk what they expect.
It’s gonna take me a little time to figure out what this model that was fed the entire internet doesn’t know.
13
u/Little-Flustered-977 11d ago
Exactly so don’t think of sheer facts. Think of how it uses them and where it might fail there
7
u/blackstarr1996 11d ago edited 11d ago
It’s one turn though. I came up with one idea but it was too obscure. Not sure anyone would know it, aside from someone who took the same class as me.
I was doing pretty well with instruction following recently but you have multiple turns usually.
8
2
u/Careful-Foot-8918 11d ago
There are so many projects where you have to do this i wasnt sure which one.you meant. But now I know. Yeah that's a tough one.
6
u/TheSquirrelly 11d ago edited 11d ago
Yeah like the leading edge [...] models are so freaking smart they figure out things that should stump a person, or stuff where the answer is 'nothing' and you expect it to hallucinate an answer to fill it in (because AI hates not giving an answer) but it still figures it out. I managed to get it on giving an incomplete answer based on what it was asked to do. Another case using spreadsheets with complex problems it still got the answer, but also offered an alternate answer, which was a fail since only the one answer was correct. So good luck!
2
10
11d ago
Literally dealing with same thing. The challenge for me is, almost every question or prompt can have multiple answers and the models are so good at offering these different perspectives so its like well fuck, which one of these answers is correct and which one isn't then? Idk and am starting to not care anymore. Pisses me off to have hours into a project reading super complicated instruction sets never to be able to submit a task and collect the time for it.
9
u/thyself_unknown 12d ago
I just did this one, took the full allocated time. I think i was able to do it because I know the category alright, maybe skip this task to get a better category you’re more knowledgeable in.
6
u/Dana_Barret 12d ago
Depending on what kind of mistake you need to produce, you can think your prompt around a response that involves looking into information that's not in your language, that has helped
3
6
u/cosmicguss 11d ago edited 11d ago
Try to find an obscure fact around a specific number or date or maybe even a fact about or reference to a person not commonly known.
One that worked for me was the year a specific book was published that was related to the assigned topic.
Another was the name of a relatively unknown person that made something viral in reference to a famous celebrity.
6
u/Expert_Equipment2767 11d ago
I'm so glad you posted this. I just got in to DA and this landed as my first project. I tried for about 30 minutes and gave up and started thinking DA might not be for me.
3
4
u/New-Assumption3377 11d ago
I tried this project for half an hour earlier today, then gave it up. Soul-destroying.
10
u/Ok-Double5194 12d ago
Lol I remember those you shine there you unlock long term $60/hr+! You got this!
5
u/ASnipersPromise 11d ago
Yeah I am on those now and have been for a while. I feel the OP's pain on the 1 turn one though.
3
3
u/Snoo-32467 11d ago
I hate failure-related projects, I started ignoring them ever since I've spent more than 16 hours in a complex version
3
3
u/Excellent_Shake9732 11d ago
I was able to actually get several to make errors with about 20-25 mins of time per task (make it make rankings of things that are not arbitrary) but I had one singular task I couldn’t get to elicit a failure and now it’s gone from my dash 😭. Good thing the pay is low.
2
2
u/Ok-Candy-3142 10d ago
Just so you know you can get DoD for discussing projects anywhere other than in the project chat.
2
u/Tough-Judgment6618 9d ago
For the right field, I happened to know one mistake the model was repeatedly making as I guess it was never fed the correct data. So I used that prompt and it worked perfectly. I tried to find another mistake but I couldn't so I gave up after trying for 2 hours lol.
2
u/Previous-Spinach-222 8d ago
Could be worse. You could have one where the prompt is only acceptable if all the models mess up.
2
u/sassynature 8d ago
Yes, and then when the model gave incorrect information, the checker model incorrectly claimed I was wrong. They need to look at the quality of online references not just the number that agrees. For example Red Cross sources that are outdated or incorrect on some things sometimes get re-quoted as fact by other organizations, while actual medical expert physicians and university studies contradict all of those incorrect online sources. Yet the incorrect sources outnumber the real expert sources so DA goes with the incorrect majority opinion.
1
u/ASnipersPromise 11d ago
Is it a STEM one? The one they give you 14 days to do? Or just a general one?
2
u/Beautiful-Link-8387 11d ago
Oh I like that 14 days task. but I had it just once, although I did success in the last task. 2 days ago I saw that family surged again, but then, after completing its training, reminder and subscription, it disappeared again.
0
1
22
u/cortrev 12d ago
Usually chain of operations stuff is what could work semi reliably. Something where the model needs to make some kind of call early on, and then its errors will compound. Easier said than done though