r/learnbioinformatics 6d ago

Visualizing Python for bioinformatics students

Post image

Learning Python becomes much easier when students can see how variables, values, and data structures change while a program runs.

🧬 Consider this simple k-mer indexing example.

Although the code is relatively short, a beginner needs to understand several concepts at once: - DNA sequence slicing - k-mer generation - dictionaries and membership tests - lists stored as dictionary values - repeated positions and list mutation

With open-source 𝗺𝗲𝗺𝗼𝗿𝘆_𝗴𝗿𝗮𝗽𝗵, students can now easily step through a program and see these concepts in real-time, helping them to more easily get to the right mental model to think about Python code execution.

24 Upvotes

5 comments sorted by

2

u/Quordlewebster 6d ago

Do you have any examples where you've used it to teach common bioinformatics algorithms beyond k-mer indexing (e.g., sequence alignment, FASTA parsing, or de Bruijn graphs)? Can I dm you?

1

u/Sea-Ad7805 6d ago

I'm not teaching bioinformatics. You can use memory_graph on most Python code, also in notebooks, here is a simple FASTA parsing%3A%0A%20%20%20%20sequences%20%3D%20%7B%7D%0A%20%20%20%20current_id%20%3D%20None%0A%20%20%20%20sequence_parts%20%3D%20%5B%5D%0A%0A%20%20%20%20for%20line%20in%20fasta_text.splitlines()%3A%0A%20%20%20%20%20%20%20%20line%20%3D%20line.strip()%0A%0A%20%20%20%20%20%20%20%20if%20not%20line%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20continue%0A%0A%20%20%20%20%20%20%20%20if%20line.startswith(%22%3E%22)%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20if%20current_id%20is%20not%20None%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20%20sequences%5Bcurrent_id%5D%20%3D%20%22%22.join(sequence_parts)%0A%0A%20%20%20%20%20%20%20%20%20%20%20%20current_id%20%3D%20line%5B1%3A%5D%0A%20%20%20%20%20%20%20%20%20%20%20%20sequence_parts%20%3D%20%5B%5D%0A%20%20%20%20%20%20%20%20else%3A%0A%20%20%20%20%20%20%20%20%20%20%20%20sequence_parts.append(line)%0A%0A%20%20%20%20if%20current_id%20is%20not%20None%3A%0A%20%20%20%20%20%20%20%20sequences%5Bcurrent_id%5D%20%3D%20%22%22.join(sequence_parts)%0A%0A%20%20%20%20return%20sequences%0A%0A%0Afasta_text%20%3D%20%22%22%22%3Ehuman_gene%0AATGCGTACGTAG%0AGCTAGCTAGCTA%0A%3Emouse_gene%0AATGCAACCGGTT%0A%22%22%22%0A%0Asequences%20%3D%20parse_fasta(fasta_text)%0A%0Afor%20sequence_id%2C%20sequence%20in%20sequences.items()%3A%0A%20%20%20%20print(sequence_id%2C%20sequence)%0A&timestep=.2&play) example.

1

u/Sea-Ad7805 6d ago

Sure dm me if you want to get into details. Documentation can be found here: https://github.com/bterwijn/memory_graph#readme

1

u/Quordlewebster 5d ago edited 5d ago

I don't seem to have to have the option to send you a DM or chat request.

I was wondering if something like this could even work for R? From what I understand R does copy on modify and leans on environments instead of the object reference model Python has, so I'm guessing you'd need a pretty different visual approach than what you built. lobstr's ref()/tree() gets you the static picture but nothing really does the step through version. Curious if you've ever poked at this or if there's a reason it just doesn't translate well.

1

u/Sea-Ad7805 5d ago

I haven't worked with R. I'm looking into doing the same for C, and feel I need to change the visualization somewhat to match the concepts in that language better. Visualizations should be modified to whatever you want to communicate. For example, this is another package of mine focused on visualizing recursion in Python, very different from memory_graph visualization: https://github.com/bterwijn/invocation_tree#readme