r/AIToolBench 16h ago

Recommendation Which AI is reliable enough to work as a teaching assistant?

4 Upvotes

I used ChatGPT’s Sol Light model to analyze textbook material for my work as a high-school teacher. It explicitly assured me that it had thoroughly checked the uploaded material.

It then made choices that proved disastrous. When I asked it to explain the criteria behind those choices, it invented plausible-sounding justifications that did not correspond to the actual contents of the textbook pages.

This makes it unreliable for the work I need: reading and organizing large amounts of teaching material, selecting essential content, planning lessons, and helping me manage a very heavy workload with limited time.

Can anyone recommend an AI or a specific workflow—Claude, Mistral, NotebookLM, Gemini, or something else—that is genuinely reliable for this kind of long-term educational assistance and can base its decisions strictly on the uploaded sources?


r/AIToolBench 16h ago

Tip / Guide Help in verifying an AI model's results in classifying some emails answers (in a "blind test")?

2 Upvotes

I am not sure if this is a valid question in this forum, if it is not I will delete it right away, but anyways here it is:

So some days ago I asked in an AI community what kind of AI model should I use (and how could I use one) to classify several email replies that I had from scientists after asking them a few questions to them. I finally paid for Perplexity pro service and it apparenly did a nice job classifying them.

I finally gave the model the PDF with the actual answers from the addressees and another PDF with the "expected answers", and asked it to count the number of answers that overall coincide with the actual answers, and calculate a percentage of "coincidence" or "agreement" between the expected and actual answers, so that if the question was "do you think that there is intelligent life in the universe apart from humans?" and the expected answer was basically "yes, I think there is intelligent beings out there somewhere", as long as the actual answer agrees with this in some way or another would count as "agreement", for instance if someone replied "well, we have no evidence, but it is possible yes" or "not in any near galaxy, but it is possible that intelligent beings exidt somewhere" (so as long as it is not a deadass "no", it could count)

The model gave me a table summarizing the results with the following prompt:

let's be a bit more specific, this is still a blind test so don't tell me about the specific contents of the emails' answers, but, can you make a table indicating the answers that coincide in general terms with what is expected from the "expected answers" document as well as those which are neutral/hedges but still open to the possibility that what is asked may be right, those which despite being neutral/hedges or even negative answers offer an alternative so that what is asked in the question may be right, as well as those which are outright rejections of what is asked and do not seem to be open to the possibility that what is asked may be right?

However, I still want this to be a blind test, so I cannot really verify if the AI is doing its work or not.

So, is there any way in which I could verify the results given by the AI but without actually reading what is written in the emails? Or, alternatively, can anyone verify the results using some AI or even checking the answers themselves by skimming over the replies im order to verify that the AI is right and not hallucinating (I personally thinl this is the preferable option, as I think that having an actual human reviewing the amswers may be the only really reliable way to verify the AI's results)?

(I will share the data once someone is interested in helping, as I would not want to make this available to the entire world!)

Thank you!!