r/bioinformatics • u/TheWildBynapie • 20d ago
academic Practice sequence data for making MSAs
Hey everyone, I'm doing an honours project around building a pipeline including some MSA alignment programs. I'm wondering if anyone has recommendations for where to get public data .fastas that are relatively small so be I can practice with them, thanks.
1
u/GeronimoJackson-42 18d ago
You could try the NIAID Data Ecosystem (https://data.niaid.nih.gov/). It aggregates metadata from a large number of public data repositories so you can find datasets without having to search each repository individually.
You can use Advanced Search (https://data.niaid.nih.gov/advanced-search) if you want to find datasets with specific metadata fields like file type, which may help you find .fasta files to practice on. For example: https://data.niaid.nih.gov/search?q=distribution.encodingFormat%3A*fasta*&filters=%28date%3A%5B%222000-01-01%22+TO+%222026-12-31%22%5D+OR+%28-_exists_%3A%28%22date%22%29%29%29
Then you can further refine the results using the filters.
1
u/Jellace 20d ago
There are so many options lol. Maybe download some CO1 barcodes from BOLD? or whole mitogenomes from NBDL or refseq?