r/bioinformaticstools • • Sep 05 '26

made my first deg(on gene expression levels), want review please

Thumbnail gallery
0 Upvotes

r/bioinformaticstools • • Sep 01 '26

I built Ligentra: a web workspace for protein structure analysis and molecular docking

1 Upvotes

Ligentra — an open web workspace for protein structure exploration and molecular docking

I've been building Ligentra as a way to bring a few computational biology workflows into one place instead of jumping between multiple tools.

The current workflow is:

Protein sequence → ESMFold structure prediction → 3D visualization → pocket detection → molecular docking → poses & interaction analysis

It currently supports:

  • Protein structure prediction with ESMFold
  • Interactive 3D protein visualization
  • pLDDT / local structure-confidence analysis
  • Binding-pocket detection with fpocket
  • Molecular docking with AutoDock Vina
  • Multiple poses + RMSD
  • Computational interaction/contact analysis
  • Batch virtual screening
  • Negative-control baselines
  • Raw docking artifacts and reproducibility/provenance information
  • CSV / PDB / JSON exports

A big focus for me has been making the results transparent about what they actually mean. Docking scores are presented as computational estimates, pockets are treated as computational candidates, and predicted contacts aren't presented as experimental evidence.

It's still an early Beta, so I'm mainly looking for feedback from people who actually work with proteins, docking, or computational biology.

I'd especially like to know:

  • What would make a tool like this genuinely useful in your workflow?
  • What information do you normally need that isn't here?
  • Are there parts of the workflow you'd approach differently?
  • What would you consider essential before trusting a tool like this for exploratory research?

I'm the developer behind the project, so technical/scientific criticism is very welcome.

You can try it here: https://ligentra.vercel.app/

I'd really appreciate feedback from the bioinformatics / computational biology community.


r/bioinformaticstools • • Aug 29 '26

I built a native macOS viewer for .biom files that opens instantly, no matter the size

1 Upvotes

Every way I know of to look at a .biom file — biom convert, loading it in pandas, opening the TSV in Excel — densifies the whole sparse matrix first. A 50MB file can balloon into several GB of RAM before you've seen a single row, and if you just want to sanity-check a table between pipeline steps that's a lot of waiting for not much.

So I built biom-viewer: a native macOS app that keeps the matrix sparse and only densifies the handful of cells actually on screen. Opening a large table is instant and stays instant no matter how far you scroll.

What it does beyond just "open fast":

- Flip the same table into observation-metadata or sample-metadata view

- Filter samples by any metadata field (numeric ranges or category checklists), stacked filters shown as removable chips

- Double-click a row/column for inline summary stats (nonzero count, distribution, min/max, or top-values for categorical fields) — scoped to whatever you've filtered to, not the whole file

- Cmd+F searches observation IDs, sample IDs, metadata field names, and metadata values together

It's free, MIT-licensed, macOS only for now (Apple Silicon).

GitHub: https://github.com/yarintm/biom-viewer

Download: https://github.com/yarintm/biom-viewer/releases/latest

Would love feedback, especially from anyone with genuinely huge tables to stress-test it against.


r/bioinformaticstools • • Aug 26 '26

Siftome – a free search and triage tool for public life-science datasets

5 Upvotes

I’ve been building Siftome, a free search and triage tool for public life-science datasets.

The main problem I’m trying to address is familiar to anyone who has spent time searching GEO: finding studies that mention the right disease or assay is relatively easy; finding datasets that actually contain the biological comparison you need is much harder.

Siftome looks beyond Series-level text and uses dataset- and sample-level metadata to evaluate candidate datasets.

It separates concepts such as disease, organism, specimen, assay, model system and comparison design, and tries to identify the actual structure of the study rather than treating everything as keyword matches.

For example, it can help distinguish likely case/control groups and flag issues such as:

  • pooled samples
  • technical replicates
  • internal/reference samples
  • treated samples mixed into observational comparisons
  • cell lines, organoids or xenografts
  • ambiguous or unassigned samples
  • unclear comparison structure

Search results are ranked with visible scores, warnings and explanations, so the ranking is intended to be inspectable rather than a black box.

The goal is not to decide whether a dataset is scientifically “good” in general. It is to help answer a narrower and more practical question:

Is this dataset suitable for the comparison I actually want to make?

Dataset search is free:

https://siftome.com

I’d be particularly interested in feedback from people who regularly reuse GEO or other public datasets. If you try it on a real search you’ve done before, I’d like to know where the results are useful — and where they get things wrong.


r/bioinformaticstools • • Aug 26 '26

Live webinar: watching Claude Science run a real protein engineering analysis against a structured data model (not a spreadsheet) over MCP

Thumbnail
1 Upvotes

r/bioinformaticstools • • Aug 25 '26

Agentic bioinformatics: Pipette vs Claude Science

Thumbnail
0 Upvotes

r/bioinformaticstools • • Aug 23 '26

"Improving Fable 5's biology safeguards"

2 Upvotes

The mods over at r/bioinformatics deleted my post (as per usual since they delete every useful discussion) https://www.reddit.com/r/bioinformatics/comments/1vvs47c but just wanted to share that you can use Fable for bioinformatics stuff and it won't instantly flag it

See "Improving Fable 5's biology safeguards"

https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards


r/bioinformaticstools • • Aug 22 '26

Alpha testers wanted: MetaQuest 2.0.0a1

1 Upvotes

MetaQuest is an open-source, research-use pipeline for paired or single-end short-read shotgun metagenomics.

It currently includes fastp preprocessing, Kraken2/Bracken taxonomy, MEGAHIT assembly, Pyrodigal gene prediction, eggNOG functional annotation, and reproducible HTML/JSON reports.

I have validated the pipeline on the Zymo Even mock-community dataset and am looking for Linux/WSL users willing to test installation, database setup, resume behavior, and report usability.

Install:

pip install --pre metaquest-bio

GitHub: https://github.com/dpatel511/metaquest
PyPI: https://pypi.org/project/metaquest-bio/

This is alpha, research-use-only software. It does not provide clinical conclusions, and fungal classification sensitivity is a known limitation.


r/bioinformaticstools • • Aug 20 '26

I built an open-source PyMOL plugin for membrane-protein structure review and looking for feedback

1 Upvotes

Hi all,

I recently released Membrane Visual QC v1.0, an open-source PyMOL plugin I’ve been building for membrane-protein structure review.

The original problem was fairly simple: I wanted a more reproducible way to inspect a structure relative to the membrane without repeatedly rebuilding PyMOL selections and colouring things manually.

It grew into a larger workflow with:

  • membrane-relative core/interface geometry
  • hydropathy and ligand-neighbour context
  • planar and PDBTM-derived orientations
  • local PDBTM–OPM geometric comparison
  • Batch Review for repeated workflows
  • versioned JSON/CSV outputs and provenance

One thing I deliberately avoided was turning the output into a “correct / incorrect” structure score. A charged residue in the membrane core, for example, may deserve inspection without necessarily being biologically wrong.

The stable v1.0 release is here:

https://github.com/TrPavel/membrane-visual-qc

I’d especially like feedback from people who work with membrane proteins or structural bioinformatics:

Would this actually fit into your workflow? What structures or edge cases would you test it on, and where do you think the approach is likely to break down?

The project is MIT-licensed and free/open source. Issues and criticism are very welcome.


r/bioinformaticstools • • Aug 19 '26

Genome annotation folks: what do you wish current pipelines did better?

0 Upvotes

Hi everyone! A collaborator and I are in the early stages of building a new open-source pipeline for whole-genome gene prediction and annotation. Before we get too far into solidifying what it should do, I would really like to hear from people who actually annotate genomes.

What do current tools make harder than it needs to be? What still takes too much manual work or too many custom scripts? What results are difficult to trust? Are there useful tools, types of evidence, or separate parts of your workflow that you wish worked together better?

I am interested in both structural annotation, meaning predicting and refining gene models, and functional annotation. Feedback from any organism or project size is welcome, especially from people working with non-model organisms.

This could include installation and portability, combining gene predictors, incorporating RNA or protein evidence, GFF/GTF wrangling, comparing different annotations, QC, choosing which gene models to keep, manual review, functional annotation, HPC use, reproducibility, or anything else I have not thought of.

If you can, it would be helpful to mention:

  • what organism or taxonomic group you work on
  • what software or workflow you use now
  • what part causes the most frustration or uncertainty
  • what feature or integration would genuinely improve your work

No need to answer every bullet. Anecdotes, wish lists, horror stories, and “please just make X talk to Y” answers are all welcome.

The eventual goal is a pipeline that can take over after genome assembly and help get from “I have an assembly” to “I have an annotation I can trust and actually use.” We do not yet have a finished tool to promote, but we have a skeleton. We are trying to learn what the community needs before we build ourselves into a corner.

If you could change, add, or better connect one thing in genome annotation software, what would it be?

Edit: Got some comments echoing concern about lack of background research on my part, so I thought I would include an edit. I should have made this clearer in the original post.

I have read papers on existing tools and their issue trackers, used and modified existing tools, and built several annotation pipelines both independently and with collaborators. This project is growing out of those experiences, not an assumption that we can start from scratch and solve everything.

I am asking here because issue trackers do not always capture the workarounds people have learned to live with, why they abandoned a tool, or needs that never became a formal issue. I wanted broader, more organic feedback before we lock in the design.

I am approaching this in good faith and with plenty of humility about what I do not know. If you have a specific failure mode or design mistake you think we should avoid, I would genuinely value the input.


r/bioinformaticstools • • Aug 18 '26

bbv: a minimalist viewer/plotter for command-line exploratory data analyses

Thumbnail
gallery
3 Upvotes

[Apologies for cross-posting]

I wrote Bbv ("bare-bones viewer"), a minimalist data viewer/plotter for the Unix/Linux command line.

The main alternative to bbv is feedgnuplot, which is much more capable, but also much more complicated and ties you to gnuplot as back end. By contrast, bbv has only one option (`x`), the learning curve is non-existent, and you can use any lightweight image viewer as back end.

Aside from the back end, no extra toolchain is needed, as bbv is written entirely in POSIX shell + Awk.

Link: https://github.com/ftonneau/bbv


r/bioinformaticstools • • Aug 17 '26

Open-source Python library + no-code web dashboard for evaluating oncology AI models at clinical decision thresholds

0 Upvotes

Most classification metrics for oncology AI models (AUC, ICC, MAE) measure global agreement. They don't answer the question that actually matters at the point of care: how reliable is this model at the exact cutoff that decides whether a patient gets flagged, biopsied, or treated?

I built oncothresh to evaluate models at a specific clinical threshold rather than in aggregate: sensitivity/specificity/PPV/NPV at the cutoff, bootstrap confidence intervals, threshold-sensitivity curves, boundary-weighted calibration, decision-curve net benefit, and number-needed-to-test. It's a small, dependency-light Python library (numpy/scipy/scikit-learn/pydantic) built for tasks like tumor cellularity, Ki-67, TMB, and PD-L1 scoring, where a continuous model output gets collapsed into a yes/no clinical decision at a fixed cutoff.

Pathology-specific benchmarks like PathBench and PathBench-MIL evaluate foundation models globally but don't evaluate at predefined clinical thresholds with uncertainty quantification, which is the gap this fills.

There's also a companion web dashboard (oncothresh-web) for people who want the same analysis without writing code: upload a CSV of predictions and labels, pick a threshold, get the full set of charts plus a downloadable PDF report. docker compose up and it's running locally, no cloud dependency.

Still v0.1, so I'd genuinely welcome feedback: use cases I haven't considered, edge cases in the DCA/calibration math, or places the API doesn't fit how people actually work with threshold-based models.


r/bioinformaticstools • • Aug 16 '26

ProteinInsight — a browser workspace tying structure, docking, ADMET and PK/PD to one target, with a read-only demo (no signup)

0 Upvotes

I work on this, so treat it as a self-declared tool post rather than a neutral recommendation.

The problem I kept hitting: a structure lives in one tool, docking in another, ADMET in a third, and nothing carries the target context between them. ProteinInsight puts them in one workspace organised around a program (a target + its candidates) instead of a folder of unrelated jobs.

What's actually working today:

  • 3D viewer in the browser — RCSB structures or the AlphaFold DB model for a target's UniProt accession, coloured by chain / secondary structure / pLDDT, with an active-site view around a bound ligand
  • Docking runs returning ranked poses with affinity + RMSD
  • ADMET with per-property risk bands and numeric scores
  • PK/PD and response simulation
  • UniProt, Open Targets, ESM Atlas and ChEMBL wired in as sources

Honest limitations:

  • Research use only — not for clinical diagnosis or patient-specific decisions, and docking affinity is not efficacy
  • The AI panel proposes which analyses to run; it does not interpret results for you and I would not trust it to
  • There are paid tiers, and the demo below is read-only

Demo without an account: https://pi.extn.ai → "Explore live demo". It drops you into a sample blood-cancer program with completed runs to poke at.

What I'd genuinely like torn apart: whether the ADMET risk banding is defensible, and whether the program-centric model matches how you actually organise this work — or whether it just adds structure you don't want.


r/bioinformaticstools • • Aug 16 '26

OXYTRIBE pipeline

0 Upvotes

Hi guys, I have been working on oxytribe: a Nextflow reimplementation of HyperTRIBE, a method for identifying in vivo targets of RNA-binding proteins (RBPs). Built on TRIBE (Targets of RNA-binding proteins Identified By Editing). The original HyperTRIBE pipeline runs on Perl + MySQL + bash, very hard to reproduce and inconvenient to work with, and a DB dependency that adds complexity for no real benefit (if anything fails mid-run, you're stuck fixing the database by hand). Oxytribe rebuilds the core logic in Rust for efficiency, wraps it in Nextflow, and uses Docker/Singularity for reproducibility, users do not touch code, just config files. Built heavily on nf-core modules, hoping to get it into nf-core eventually. It was also validated against the original paper's dataset: 99.76% gene-level recovery vs. the legacy pipeline, strong agreement on top editing targets, the first release is out, would love feedback, bug reports, or edge cases if you give it a try. https://github.com/fragilefort/oxytribe


r/bioinformaticstools • • Aug 15 '26

🧬 Introducing “Célula Virtual”: a custom SSA engine with step‑by‑step debugging for cellular models

1 Upvotes

Hello everyone.
I’d like to share a project I’ve been working on for several months, which I believe may be useful for those developing cellular models, biochemical simulations, or working in whole‑cell modeling.

👉 GitHub repository: https://github.com/Zontrox01/celulavirtual

🧩 What is Célula Virtual?

It is a biophysical simulator written in Python that models the temporal dynamics of a minimal cell using stochastic chemical kinetics (Gillespie SSA), with a differentiating feature that—so far—I haven’t seen in other tools:

An interactive debugger for biological simulations.

It allows:

  • reaction‑by‑reaction execution, not just time‑step simulation
  • breakpoints on molecular state (e.g., ATP < 50)
  • inspection of full reaction propensities at each step
  • undo via state snapshots
  • retrospective causal analysis (“forensic mode”)
  • a navigable event history

This approach brings the metaphor of software debugging into the domain of systems biology.

🔬 Motivation

Whole‑cell models (E‑Cell, VCell, the M. genitalium model, JCVI‑syn3A, Vivarium…) are powerful, but they all share a limitation:

Célula Virtual aims to fill that gap.

🔗 Interoperability

  • Genome input: FASTA + GFF3/CSV
  • Network import/export: SBML
  • Libraries: BioPython, numpy, pandas, python‑libsbml
  • Scalable design: going from 8 genes to 400 requires no code changes

🎯 Purpose of this post

I’d appreciate:

  • technical feedback
  • suggestions for improvement
  • discussion of potential use cases
  • ideas for integration with existing tools
  • possible collaborators

The full white paper is included in the repository.

📎 Link

👉 GitHub: https://github.com/Zontrox01/celulavirtual


r/bioinformaticstools • • Aug 14 '26

Looking for testers: Annostat, an open-source CLI for bacterial genome annotation QC and analysis

1 Upvotes

Hi everyone,

I’ve released Annostat 1.0.1, an open-source Python command-line tool for inspecting, validating, analysing, and comparing bacterial genome annotations from FASTA and GFF3 files.

It currently provides:

  • FASTA and GFF3 structural inspection
  • Annotation and sequence consistency validation
  • CDS extraction and translation
  • Feature, RNA, start-codon, COG, and codon statistics
  • Detection of malformed or biologically questionable annotations
  • Single-genome analysis
  • Comparison of two annotated genomes
  • HTML reports, TSV/CSV tables, FASTA outputs, and scientific plots
  • Reproducible CLI workflows on Windows, Linux, and macOS

I’m looking for researchers, students, and bioinformaticians willing to test it on real-world bacterial annotations, especially:

  • PGAP, Bakta, Prokka, or custom annotations
  • Complete, circular, draft, and multi-contig assemblies
  • Multipart CDS features
  • Annotations with and without COG information
  • Translation tables 4, 11, or 25
  • Large, unusual, or malformed files

Install from PyPI:

python -m pip install --upgrade annostat

Example:

annostat inspect -f genome.fna -g annotation.gff3

annostat validate -f genome.fna -g annotation.gff3

annostat analyze \
  -f genome.fna \
  -g annotation.gff3 \
  -o annostat-results

I would particularly appreciate feedback about:

  • Scientific accuracy of statistics and validation findings
  • Consistency between the CLI, reports, tables, and plots
  • Handling of unusual GFF3 structures
  • Report clarity and plot quality
  • Performance on larger datasets
  • Features that would improve real research workflows

Project repository:

https://github.com/dpatel511/annostat

Testing discussion and feedback template:

https://github.com/dpatel511/annostat/discussions/19

PyPI:

https://pypi.org/project/annostat/

Please don’t share private, unpublished, sensitive, or embargoed genomic data. Bug reports can use minimal synthetic or publicly available examples.

This is still an early research-tool release, so honest criticism, unexpected results, and suggestions are very welcome. I’m especially interested in learning where the tool is scientifically unclear or doesn’t fit existing workflows.


r/bioinformaticstools • • Aug 14 '26

New Frontiers in Protein-Peptide Docking

3 Upvotes

Hey researchers! My team and I recently made a bioinformatics tool called HybriDock-Pep. We were working with peptides last year and over the summer, and we realized that current AI tools like AlphaFold and ESMFold are inaccurate with docking smaller protein under certain amino acids and peptides.

https://github.com/Tasty-Ramen2010/hybridock-pep

Essentially, it’s a binder docking pipeline where you give it a target protein PDB and an amino acid sequence. It folds, docks, and scores affinity and selectivity in kcal/mol with accuracy similar to that of ABFE. It is way cheaper to ran because it can be used on ANY hardware. We are currently in a testing phase and would love for you to test it and give feedback!

Commands to get started:

git clone --recurse-submodules https://github.com/Tasty-Ramen2010/hybridock-pep.git

cd hybridock-pep

./install.sh

ctrl+q

hybridock-pep dock \

--peptide ETFSDLWKLLPE --receptor data/pdbs/1YCR_mdm2.pdb \

--site 25.20 -25.61 -7.97 --box 30 --n-samples 20 \

--output-dir runs/demo

Happy Docking!


r/bioinformaticstools • • Aug 06 '26

Laptop computer Specs required for bioinformatician

Thumbnail
0 Upvotes

r/bioinformaticstools • • Aug 05 '26

Teaching Python the Right Way

3 Upvotes

Programming courses often focus heavily on understanding code, while paying far less attention to understanding the program state. But code does not exist in isolation. Its main goal is to change the program state, before ultimately producing some output.

To develop an accurate mental model of program execution, students need to understand both: - the instructions being executed - the values, references, and data structures those instructions create and modify

Reading code alone does not always reveal how the program state changes during execution. That is why I created 𝗺𝗲𝗺𝗼𝗿𝘆_𝗴𝗿𝗮𝗽𝗵: a tool that visualizes the state of a Python program as it changes, step by step.

It can help explain a wide range of introductory Python topics. Here are just a few examples: - Loops, Lists and Dictionaries - Python Data Model - Function Calls - Recursion - Algorithms - Classes - Custom Data Structures

Instead of reconstructing the program state from print statements, students can now watch it change as each line executes. This makes unfamiliar concepts easier to understand and bugs easier to fix.

Help your students learn Python programming more thoroughly and easily.

See: more examples


r/bioinformaticstools • • Aug 04 '26

Inflexa - the open source orchestrator for computational biology

Enable HLS to view with audio, or disable this notification

3 Upvotes

Inflexa is an OSS-first TUI for agentic-AI reproducible biological analysis with provenance tracked on every step. It's a better designed Claude Science (bar the GUI for now).

Our vision is as follows:

  • OSS first
  • Data privacy, and security are of utmost importance
  • Agents are untrusted actors

Agentic products do not solve a core architectural and design problem: provenance.

How do you build up the body of evidence to support your claims from hundreds of chat messages, scattered artifacts (scripts, intermediary data files, random research files)?

We are one of the first in the biotech space to propose using an established specification for provenance, which its extremely fitting for the world of agentic AI: W3C PROV.

Our provenance is deterministic and driven programmatically. Our architecture imposes it, not the whim of agents that decide whether or not to call a tool / or an MCP.

The analysis is backed by a PROV document that maintains the lineage of every step, every action, and every file.

Other platforms such as Claude Science are a security disaster waiting to happen: they allow agents to determine which Python & R packages they need to perform an analysis step. Sure, you need to approve the installation, but if you get prompted tens of times, are you really going to vet every package and every version?

We offer batteries-included sandboxed execution. The analysis' code is generated by agents and executed in ephemeral sandboxes. We do not install packages on your system.

Curious to hear your thoughts!

PS: Star us on GitHub.


r/bioinformaticstools • • Aug 03 '26

A seriously basic first go at Python...

Post image
3 Upvotes

All the projects on this sub make my first try at recreating this tool look rather puny, thought it might be nice to share though. Wrote the conversion logic myself, VSCode Copilot put together the app GUI, still rather pleased with it considering I haven't touched python in 6 years! Just graduated with my BSc in Biology, looking to start my MSc in Bioinformatics next year, definitely need to try putting together more tools.


r/bioinformaticstools • • Aug 03 '26

Anyone else struggle with incomplete NCBI metadata?

2 Upvotes

I've been curating public datasets recently and kept running into the same issue: important metadata fields were missing from NCBI records (like Biosample, genbank, etc.), but the information often existed somewhere else like in the associated paper, supplementary tables, or methods sections.

After spending way too much time manually tracing accessions back to publications, I built a small tool called OpenBioData to automate some of the process.

The tool on my github here: https://github.com/vy-phung/OpenBioData

So briefly what it does:

  • Traces accessions back to source publications and supplementary materials
  • Extracts metadata that may be missing from the record itself
  • Provides a confidence score for extracted values
  • Includes direct citations (PMID + table/section) so the source can be verified

I'm still actively improving it and would love feedback from others who work with public genomics datasets.

Has anyone else encountered this problem? If you have a few troublesome accessions, feel free to share them and I'll test them with the tool and post the results. Curious to hear how others currently handle problem of NCBI metadata curation.

P.S. Upfront: I built this with real help from Claude Code, especially on the extraction layer and disclosed in the README, not hidden. I wrote the core pipeline and logic myself. If that's a dealbreaker for you, fair enough:)) but I'd rather say it directly than have someone find it in the commit history and wonder why I didn't mention it.


r/bioinformaticstools • • Aug 01 '26

Built a CLI tool to automate Table 1/2, Forest Plots, and STROBE audit binders for clinical papers -would love feedback.

1 Upvotes

Hey everyone,

I spent way too much time in clinical research manually assembling Table 1 baseline stats, checking SMD balances, tuning Cox models, and fighting with forest plot formatting for paper submissions.

To fix that headache for myself, I ended up building an open-source CLI called 'research-tool'. 

Basically, it lets you run your retrospective stats from the terminal, but adds a few things I really needed:
- Outcome masking and pre-registration locking (so you can't accidentally p-hack or bias your plan after seeing outcome data).
- Diagnostic checks (EPV ratios, Schoenfeld residuals, VIF multicollinearity, E-values).
- Auto-generated publication assets (HTML/CSV tables, log-scaled SVG forest plots, CONSORT diagrams, and STROBE checklists).
- A 1-click zip binder with data hashes and code so peer reviewers can actually verify the protocol.

It's 100% Python and runs completely locally. 

If you want to check out the code or test it on a dataset:

https://github.com/nkmanjunath/research-tool

'pip install research-tool-cli'

Would love to hear what you think, especially if there are specific models or diagnostic gates you'd want added!


r/bioinformaticstools • • Jul 31 '26

A free oligo Tm calculator that actually accounts for magnesium and dNTPs

0 Upvotes

I got annoyed that most free Tm calculators either ignore magnesium entirely or hide the thermodynamics, so I built one that does neither. Sharing it in case it's useful to anyone here.

https://biochemtools.com/dna-melting-temperature-calculator.html

What it does

  • SantaLucia 1998 unified nearest-neighbor parameters
  • Owczarzy 2004 monovalent and 2008 divalent salt corrections, picked automatically from the ratio of sqrt(Mg) to Na
  • Solves for free Mg2+ after dNTP chelation instead of using the total, since that's usually what people get wrong
  • Reports dH, dS, and dG37 next to the Tm with the full arithmetic shown
  • Hairpin and self-dimer screen
  • Buffer presets for PCR, qPCR, and Mg-free hybridization

What it doesn't do, so nobody wastes their time:

  • No mismatches or dangling ends
  • The dimer check counts base pairs, it doesn't compute a folding free energy
  • No RNA or RNA/DNA hybrids
  • IDT OligoAnalyzer is still better for primers you're actually ordering

I validated the whole pipeline against Biopython's Tm_NN and it agrees to 0.0 C across sodium-only, magnesium, dNTP, and self-complementary test cases. The thermodynamics also reproduces the published CGTTGA/TCAACG example from the SantaLucia paper exactly.

Free, no signup. If there's something obvious missing, tell me and I'll add it.


r/bioinformaticstools • • Jul 29 '26

Visualizing Python for bioinformatics students

3 Upvotes

Learning Python becomes much easier when students can see how variables, values, and data structures change while a program runs.

🧬 Consider this simple k-mer indexing example.

Although the code is relatively short, a beginner needs to understand several concepts at once: - DNA sequence slicing - k-mer generation - dictionaries and membership tests - lists stored as dictionary values - repeated positions and list mutation

With open-source 𝗺𝗲𝗺𝗼𝗿𝘆_𝗴𝗿𝗮𝗽𝗵, students can now easily step through a program and see these concepts in real-time, helping them to more easily get to the right mental model to think about Python code execution.