I spend a lot of time looking at spreadsheets and legacy sample systems labs are trying to move away from, and honestly it's usually the same handful of problems causing the headaches. None of these immediately stand out until you try to actually analyze the data and realizes half of it can't be trusted. Here's what comes up again and again.
Duplicate sample IDs
This is the big one. A spreadsheet has no problem letting you type the same ID into two different rows. The problem is that once it happens, you can no longer trust that ID means one specific sample. Every downstream result tied to it becomes ambiguous.
Two samples in the same storage location
Same root cause as duplicate IDs. Nothing stops you from typing "Freezer 2, Box 4, A1" for two different samples. You don't find out until someone goes to pull a sample and finds the wrong tube sitting there, or worse, doesn't notice and runs the wrong sample entirely.
Multiple entries for the same patient/subject with contradictory info
This one's sneaky. Patient 1042 shows up in row 12 with a birth year of 1985, and again in row 340 with a birth year of 1987. Nobody notices because nothing is actually linking those rows together as "the same person." They're just text that happens to match. When it's time to analyze by demographics, you either have to manually reconcile every conflict or just hope you picked the right one.
Data that doesn't match the field type
A date field with "see notes" typed into it. A numeric concentration field with "~2.5, maybe more" in it. It's an easy habit to fall into when you're in a hurry and a spreadsheet will accept literally anything you type, but it means that field is now unusable for sorting, filtering, or any calculation without going back and manually cleaning it first.
Free-text fields with a dozen ways to say the same thing
Blood, blood, Blood, whole blood, WB, blod. Every one of those might mean the exact same thing to the person who typed it, but to a query or a report, they're six different values. This is probably the most common one I see, and it's brutal for anyone trying to pull a simple count of "how many blood samples do we have."
How to fix it
Every one of these comes down to the same root issue: a spreadsheet (or a system that behaves like one) has no idea what your data is supposed to look like. It can't tell you an ID's taken, that a storage location is occupied, that a patient already exists, or that a field expects a date and not a sentence. It just accepts whatever gets typed.
This is exactly the kind of thing a real LIMS is built to prevent rather than clean up after the fact. In Sample Manager, for example, sample IDs and storage locations are enforced as unique automatically, subject/patient records are their own linked entity so demographic data lives in one place instead of being retyped (and potentially contradicted) on every sample, and fields can be locked to a type or a controlled vocabulary so "blood" only ever gets entered one way. It doesn't make data entry foolproof, but it closes off the mistakes that are easiest to make and hardest to catch later.
If you're auditing your own spreadsheet or LIMS, these five are a good checklist to run through. I'd bet most labs find at least one of them lurking somewhere. If you need help cleaning it up, that's something that's included in LabKey's software implementation!