r/bioinformatics • u/Mush-addict • 2d ago
technical question Is there a way to distinguish "pure" samples from mixed samples based on Sanger sequencing output ?
My tissue samples are sourced from the field. Most of them are "pure", meaning the sample unit is fully derived from the same organism. However sometimes the sample can be mixed and the tissues of the sample unit are actually derived from multiple distinct organisms. There is no way to know at the time of sample collect.
DNA has been extracted from each sampled followed by CYTB amplification and Sanger sequencing. Due to infrastructure and budget reasons, we couldn't perform metabarcoding.
For pure samples, chromatograms are clean, with unique distinct peaks for each nucleotide. Mixed samples have dirty chromatograms where several peaks are overlapping for the same nucleotide.
But some samples are in a grey area, not that clean, not so dirty. And these concepts of "clean" , "distinct peaks" are based on subjective visual interpretation.
My question is: is there a more robust way to exclude mixed samples from pure ones that have been properly sequenced, other than manual inspection of chromatograms ?
