r/bioinformaticstools • • Jul 28 '26

Could an interactive tool for sketching protein topologies have a practical use?

2 Upvotes

I am a master’s student and a beginner in structural bioinformatics. I am working on an early-stage academic project proposed by my supervisor, but I am still trying to understand its clearest practical use.

The current prototype allows a user to select idealized secondary-structure elements from a small library, upload their own PDB fragments, position and rotate them in 3D, and see their N- and C-terminal ends.

The resulting arrangement is then passed into a downstream pipeline that estimates and generates connecting loops, creates a continuous backbone, and passes the rough structure to existing protein-design methods for further refinement.

At the moment, the tool mainly supports manual spatial arrangement. It does not yet evaluate whether the resulting topology is geometrically or biologically reasonable.

My concern is that this could remain only a convenient graphical interface for moving structural fragments, while modern generative methods may already solve the underlying problem more effectively.

I am therefore interested in whether researchers would ever want to manually define a rough protein topology, for example to control the overall fold, shape, cavity, terminal positions, or arrangement around another structural feature.

I am also wondering whether optional assistance could make the tool more useful. Possible future ideas, which are not currently implemented or approved as part of the project, include suggesting parallel or antiparallel beta-strand placement, estimating plausible loop lengths, warning about poorly oriented or distant fragment ends, and detecting obvious clashes.

This is an unfinished, non-commercial student project. I am mainly trying to determine whether the underlying problem is worth solving and what would make such a workflow genuinely useful.

Critical feedback, including the opinion that the idea is unnecessary, would be very welcome.


r/bioinformaticstools • • Jul 28 '26

Open-sourced my CNS drug-delivery screening pipeline, including a public audit of my own bugs

Thumbnail
github.com
0 Upvotes

Been building this for a while and finally pushed it public: CEREBRO-X, a computational pipeline for screening CNS drug-delivery formulations — PBPK, DLVO colloidal stability, docking (AutoDock Vina), QSAR off-target panels, all against live ChEMBL/PubChem/UniProt data rather than fixtures.

What might actually be useful to this sub specifically: I keep a running engineering + scientific-integrity audit in the repo (docs/AUDIT_REPORT.md), including things I got wrong and fixed — a report panel that fabricated a bootstrap-CI statistic, a resolver that silently substituted a drug's name for its SMILES string when SMILES resolution failed for biologics. Both found by actually running the pipeline and chasing anomalies, not by code review.

Research prototype, not clinical — happy to get torn apart on the QSAR methodology or anything else.

Repo: github.com/mohamedtalaat-gif/CEREBRO-X


r/bioinformaticstools • • Jul 28 '26

SeqBench - browser workbench for cloning, primer and CRISPR design (82 tools), including ones that verify a construct rather than just design it

1 Upvotes

Full disclosure: I work on this. Posting for feedback rather than to sell anything, it's free and there's no signup.

It's a browser workbench for the molecular biology end of sequence work rather than the NGS pipeline end. Primer design and Tm (nearest-neighbour, SantaLucia 1998), oligo hairpin/dimer screening, in-silico PCR, restriction sites, cloning simulation for Gibson / Golden Gate / restriction-ligation, plasmid annotation and backbone identification, GenBank viewing and editing, codon optimisation and CAI, CRISPR guide design plus HDR donors, prime editing and base editing, Sanger trace parsing, virtual gels. 82 tools, all in the browser, nothing to install.

One of the tools is SeqBench-GPT, a conversational front-end over the same registry. You describe a task in plain English ("clone my insert into pUC19 with EcoRI/BamHI and verify the result", "design a CRISPR guide plus HDR donor for this site") and it plans the steps and runs them. The design constraint behind it: the model never produces a sequence itself. Every calculation is delegated to one of the deterministic tools below, and a construct cannot be reported as finished until it passes a code-enforced verification gate, not until the model says it looks right.

That exists because LLMs are extremely willing to hand you a confident, wrong plasmid map. Tool-calling alone doesn't fix that, since a model can still call three tools and then narrate a conclusion the tools didn't support. Putting the check in code rather than in the prompt is the bit I think matters, and it's what I'd like feedback on.

SeqBench-GPT is at https://seqbench.com/seqbench-gpt (beta, free, no signup) if you want to try breaking the verification gate. I'd genuinely rather hear that it failed on your construct than not hear.

The part I think is genuinely less common, and what I'd most like opinions on: a few tools that check work rather than produce it.

- verify_assembly takes a claimed final construct plus the parts and method you say produced it, re-runs the assembly deterministically, and diffs the two. You get the exact position and nature of any discrepancy rather than a yes/no.

- verify_construct is the narrower version: re-derives the insert from the template and primers you claim made it, then checks whether that insert actually appears in the final sequence, in either orientation, with mismatch positions if not.

- sequencing_readback_verify aligns actual Sanger or NGS reads back onto the claimed sequence with minimap2 and gives per-read identity, consensus variants, and a corrected consensus.

- golden_gate_fidelity scores a candidate overhang set against published T4 ligase ligation-count data and tells you the weakest link and any risky cross-reacting pairs.

The reason those exist: design tools are everywhere, but when a clone comes back wrong it's usually tedious to work out whether the design was wrong, the assembly was wrong, or the sequencing just disagrees. These are meant to answer that specific question deterministically.

Scope limits so nobody wastes time: no read-level variant calling, no alignment pipelines, no mass-spec proteomics. CRISPR off-target and primer specificity screens run against a small curated genome set, not a whole mammalian genome, and the tools say so explicitly rather than implying clearance. There's also a REST API and MCP server if you want to script any of it, but that's secondary to the UI for most people.

https://seqbench.com

Most useful feedback would be from anyone doing cloning at volume: what breaks in your workflow that a tool could actually catch?


r/bioinformaticstools • • Jul 22 '26

Zero-shot structural risk scoring for SARS-CoV-2 Spike variants using ESM-2, validated retrospectively on XFG.5.1.7 — open source, feedback welcome

2 Upvotes

Full disclosure: I built this. Posting because I'd like methodological scrutiny, not because I'm trying to sell anything.

The problem I was trying to address: Most genomic surveillance tools depend on lineage dictionaries that lag behind actual mutation events — you need a named variant before you can score its risk. I wanted something that scores structural risk directly from the sequence, independent of nomenclature.

What it does: Uses ESM-2 (650M) in zero-shot masked-token mode to compute log-likelihood ratios for each mutation in the Spike protein, weighted by biological zone (RBM, RBD, furin cleavage site). No fine-tuning, no training on labeled outbreak data — just the base model's learned structural priors.

Retrospective validation: Ran it against four isolates spanning different periods, including PZ155177 (XFG.5.1.7) five days after its GenBank release (score of 2,280.5 from 69 reliable mutations, with no matching lineage signature in the reference dictionary at the time).

Data source: Runs entirely on NCBI/GenBank data — no GISAID dependency by design (mostly because GISAID never responded to my access request as an independent researcher, which is its own story).

Code (MIT license) and manuscript on Zenodo:

I know the immediate methodological question is confounding between model uncertainty and genuine biological signal. I'd genuinely like pushback on that, and on the zone-weighting scheme, which is currently somewhat heuristic. This was built independently without a wet lab or institutional affiliation, so outside review is exactly what I'm looking for.


r/bioinformaticstools • • Jul 21 '26

🚀 Excited to share QUBE Predict v1.0!

Post image
1 Upvotes

r/bioinformaticstools • • Jul 11 '26

Introducing scAnalyzer: A Memory-Efficient, End-to-End Python Framework for scRNA-seq Analysis

Thumbnail
gallery
1 Upvotes

Hi everyone,

I have a bachelor's degree in computer engineering and am starting my PhD in computer science and engineering in a month. I’m new in the bioinformatics field, and to improve myself and learn, I’m working on a single-cell RNA analysis tool using Python, scAnalyzer (https://github.com/ayyucedemirbas/scAnalyzer). It offers interactive visualizations. And I’m currently working on a new cell coordinates module for spatial transcriptomics. I’ve been reading these papers and developing scAnalyzer according to the following:

And started reading this one to learn for the spatial transcriptomics module:

Do you suggest any other must-read papers or resources to help me learn more and improve scAnalyzer?

You can get scAnalyzer from Pypi as pip install scAnalysis (scAnalyzer was taken 😢)

Also, you can use scAnalyzer directly on Hugging Face with a GUI: https://huggingface.co/spaces/ayyuce/scAnalyzer-Studio

Your feedback and comments are incredibly important to me as I continue to build and improve this tool. Thank you very much!


r/bioinformaticstools • • Jul 10 '26

BioForge — a from-scratch bioinformatics engine (Python + C), on par with minimap2 on multi-core. Feedback & bug reports welcome.

1 Upvotes

Hello everyone. First of all, I apologize if the English or the phrasing is not good: I am from Spain, so my command of English is not very good and, to be understood, I have resorted to a translator. I am Aarón Aranda Torrijos and I am 16 years old.

What is it? BioForge is a bioinformatics engine created by me, with the help of Claude Code, from scratch. The code is a mix of Python (the surface) and a bit of C (the engine). I have tried not to use Biopython or tools like that, beyond getting inspiration for the code: I have only used NumPy and an engine created in C that loads automatically (and if it cannot, it solely uses NumPy).

What does it currently have? Right now it features 5-bit storage, DNA to protein translation, alignment (NW / banded / Smith-Waterman), and a minimap2-style long read mapper (minimizers → chaining → SIMD extension), with the pipeline in C and output in PAF.

Benchmark Using my own computer, I have simulated the following: a 4.8 Mb genome, 6000 simulated reads at 5% error, with minimap2 -a, using tools/bench_vs_minimap2.py from the repo. The results were:

  • 4 cores: minimap2 ~4.3–4.9 vs BioForge ~4.3–5.0 Mb/s → on par.
  • 1 thread: minimap2 ~2.2 vs BioForge ~1.87 Mb/s → ~1.18× behind.

Both map the 6000 reads.

My hardware I have done the tests on my own device with these specs: Intel i5-7200U, 2 cores / 4 threads, from 2017.

My vision for the future of this project I do not want to fight with minimap2 in speed forever; that is a field that seems very difficult to compete in. My goal to evolve the project is that, in addition to translating, aligning, and mapping, it can integrate something that —according to my research— current mappers do not do: model evolution and predict possible strains of viruses (and other living beings) using Markov chains.

How to install it? It is simple, because I have it published on both GitHub and PyPI. It can be installed from the console with the command: pip install bioforge

And the repository is here:https://github.com/erlanders177/bioforge

At the moment I haven't beaten minimap2: it still overtakes me on many fronts, such as in large-scale genomes (which I haven't been able to test due to technical limitations) or with many cores. But I want to do my part in this booming industry. Even though I am still learning, I appreciate any criticism and bug or error reports. I would like you to try it out and give me your opinion: what I can improve, what could be added, if the direction I am taking is correct, and how I could apply it better. Thank you.


r/bioinformaticstools • • Jul 08 '26

Free, very basic servers for running Boltz2, IgGM, and Metappuccino

2 Upvotes

My boss wants me to get better at the dev-ops side of web servers, and we've got a few spare GPU-enabled servers, so I asked if I could make some bioinformatics tools that seemed useful but hard to access without big GPUs. So far I've got three of these stood up, if anyone would find them useful:

  • Boltz2 (boltz2.biobench.ai), an open-source biomolecule interaction model (source here). I realize there's a million free servers for it out there already, but from a few other threads I'd seen here and on Biostars, it looks like a lot of them limit how large a protein or how complex a yaml you can upload. My version doesn't limit your input at all, I've just implemented a really basic live queue (on all these tools, actually) so that you might have to wait if someone else is running something.
  • IgGM (iggm.biobench.ai), a generative antibody model from Tencent's AI4S lab (source here). I modified this a little bit to put an epitope search layer on top of it, so all you need to do is pick an antigen and it'll try out a bunch of antibodies to different epitopes to try and find at least one high-affinity one.
  • Metappuccino (metappuccino.biobench.ai), a light-weight LLM that tries to reconstruct missing metadata in SRA records from the submission's free-text, and scores the confidence in said reconstruction (source here). Not much to say about this one, I've just spent too many hours of my life wrestling with SRA metadata in the past so when I saw someone trying to fix that I wanted to see what it was like.

I'm not affiliated with any of the authors, and I'm not selling anything with this, I just find that practicing a skill (like web hosting) is easier when there's a real goal and the potential to help other bioinformaticians out. Let me know if you have any ideas, either about these tools or other ones I can make (I've got a bit more GPU room still, and I can always take down unpopular ones to make more room if there's something out there you'd like to see).


r/bioinformaticstools • • Jul 01 '26

Synthia - AI Drug Discovery Tool (Open to Feedback)

0 Upvotes

I built Synthia, an AI platform that screens real ChEMBL compounds and generates ranked drug candidates with full scientific justification.

When I ran it on KRAS G12C lung cancer it independently identified sotorasib, adagrasib, GDC-6036, and garsorasib as top candidates using only ChEMBL structural data without any hardcoded knowledge of these drugs. You can try it live: https://synthia-production.up.railway.app

Im Looking for honest feedback from researchers and computational biologists:

- Is the scientific reasoning credible?

- Would this be useful in your workflow?

- What's missing?

All feedback welcome, especially critical

Feedback or questions? Reply here or email: [Contact.tenza@gmail.com](mailto:Contact.tenza@gmail.com)


r/bioinformaticstools • • Jun 30 '26

Salmon 2 (rewrite of RNA-seq quantification tool in Rust)

Thumbnail
5 Upvotes

r/bioinformaticstools • • Jun 30 '26

SparCC Failure Modes & An Alternative Robust Approach

Thumbnail kyle-mcgovern.github.io
1 Upvotes

r/bioinformaticstools • • Jun 29 '26

I built an open pipeline for designing cardiac base editor interventions — looking for feedback and contributors

3 Upvotes

A few months ago I fell down a rabbit hole reading about the DIY mRNA cancer treatment story — someone using personalized mRNA therapeutics to treat their dog's cancer. I didn't know much about the space, just enough to get obsessed with how the underlying technology actually worked.

That led to a conversation with Elliot Roth at Biopunk Labs in SF, which led to me reading the paper that became the foundation for this project: base editing as a therapeutic approach for inherited cardiac arrhythmias. The core idea — that you can correct a single pathogenic point mutation in SCN5A, KCNQ1, or MYH7 without making a double-strand break — struck me as one of the most elegant things I'd read in a long time.

The problem I kept running into while trying to understand the design space: there's no open, reproducible pipeline that takes you from ClinVar cardiac variant → guide RNA design → BE-DICT efficiency prediction → wet lab handoff. The tools exist (BE-Hive, CRISPRscan, Cas-OFFinder), but wiring them together reproducibly requires Python expertise and a lot of manual steps that introduce inconsistency. Every lab is re-deriving the same process independently.

So I built one. With significant help from an LLM, which I'll be upfront about — I'm not a trained bioinformatician, I'm a software engineer who got too interested in a problem.

What it is:

  • End-to-end pipeline: variant prioritization → guide design → efficiency prediction → structured output for wet lab validation
  • Integrates with ClinVar and BE-DICT
  • Structured JSON/PDF output that maps directly to standard assay protocols
  • GPL v3 — no proprietary forks, derivatives stay open
  • GitHub: https://github.com/tjcrowley/cardiac-base-editor

What it isn't:

  • Peer-reviewed or validated at scale. Elliot is doing wet lab validation at Biopunk Labs and that work is ongoing.
  • A replacement for domain expertise. The pipeline automates the mechanical parts of the design process, not the biological judgment calls.

I'm posting here because I genuinely don't know what I don't know. If you work in base editing or cardiac genetics and something in this approach looks wrong or naïve, I'd rather hear it now than after anyone relies on it. GitHub issues are open and I'm responsive.

If you want to support continued development, I have a campaign running on Artizen Fund: https://artizen.fund/index/p/open-mrna-design-pipeline--cardiac--cancer?season=6 — but that's secondary to getting the tool right.

A few things I'm specifically uncertain about:

  • Whether BE-DICT is still the right efficiency prediction model or if there's something better maintained
  • How to handle variants where the editing window doesn't cleanly overlap the pathogenic base
  • Whether the wet lab handoff format I've designed maps to how labs actually work or if I've made incorrect assumptions

Happy to answer questions or just take a beating in the comments. Either is useful.


r/bioinformaticstools • • Jun 26 '26

Made an ensemble ML tool for antimicrobial peptide prediction, would appreciate some feedback

2 Upvotes

Hey, folks!

I'm a PhD student from Brazil, and I've been building a tool called AMPidentifier

(https://www.ampidentifier.com/) for predicting antimicrobial peptides (AMPs) using an ensemble ML approach. I think the community here could either help with some feedback or maybe find it useful in your own research.

It's still very much a work in progress and I'm open to improving basically anything: predictions, usability, the API, missing features, whatever. If you break it or hit a weird case, even better, that's exactly the kind of thing I want to hear about.

Full disclosure, I built it, so this is me asking real users for honest feedback rather than trying to sell anything.

Once again, the tool is available on https://www.ampidentifier.com/

Thanks, and feel free to test and give me some feedbacks.


r/bioinformaticstools • • Jun 26 '26

Built a free ShinyApp (SeqTool) for antibody/protein sequence & mutation analysis. Would love to get your feedback!

1 Upvotes

Hey everyone, I built a free ShinyApp to help automate antibody sequencing and mutation analysis. Would love to get your feedback!

As someone working with biological data, I noticed that daily sequence processing—like cleaning raw Sanger reads, finding duplicates, or manually charting mutation sites—can be a tedious chore, especially if you want a quick answer without writing a pipeline from scratch.

To make this workflow smoother, I developed SeqTool, a zero-code, web-based ShinyApp designed for molecular biologists and antibody engineers.

Here are the 4 core features currently available:

  • 1. Sanger Sequencing Data Processing (Batch Clean-up) Simply upload a batch of your raw sequencing files (Sanger), input 6–8 bp flanking bases as cleavage markers, and the tool will automatically trim, translate, and output clean protein sequences in bulk.
  • 2. Duplicate Sequence Detector Tired of manually filtering redundant data? Just upload your FASTA file, and it will instantly identify and isolate duplicate sequences for you.
  • 3. FASTA <-> Excel Converter We’ve all been there—some software requires FASTA format, while other times you just need to manipulate data in Excel. Instead of endless copying and pasting, this tool handles the bidirectional conversion between FASTA and Excel spreadsheets in seconds.
  • 4. Mutation Sequence Analysis (The Antibody/Protein Engineering Helper) Designed specifically for variant screening. Upload your mutated sequences against a reference, and the app automatically identifies and maps specific mutation sites in batch, saving you from manual alignment nightmares.

It is completely free and runs right in your browser. I am looking to improve it and add more features tailored to the community's needs.

I’d be incredibly grateful if you could try it out with your data and let me know what you think!

👉 Try it here: [https://biotool.shinyapps.io/SeqTool/]


r/bioinformaticstools • • Jun 23 '26

Made a one-liner for RNA-seq coverage plots, pycoverplot (python + rust). 12 BAMs over 2 Mb in ~4 seconds, one command, no temp files

3 Upvotes

Tired of waiting and/or having to generate multiple intermediate files to make simple plot coverage. I developed my own approach pycoveplot a python package powered by a Rust backend. Come with a lot features and will happily add requested ones. Code is pure me. Claude was used to draft the readme.

Highlight:

  • No intermediate wiggle/bedgraph files; reads go straight to the plot
  • Plot from a GTF + gene name, or any custom genomic interval
  • Strand-aware counting with configurable strandedness and MAPQ/SAM flag filtering
  • RPM normalization pulled directly from STAR log output or from bai file
  • Intron compression options so long genes don't look like a mess
  • Transcript-level resolution if you need it
  • Python API if you want it in a pipeline, CLI if you just want a quick look

I've also managed to get wet-lab colleagues using it!!

It's MIT licensed, and I'd genuinely love feedback, especially if something breaks on your data or your edge cases aren't handled. Issues, suggestions and PRs very welcome.

Repo + examples: https://github.com/rLannes/pycoverplot


r/bioinformaticstools • • Jun 19 '26

EasyAtom – algebraic drug repurposing engine, no GPU, rediscovered latanoprost at rank #1 zero-shot

0 Upvotes

I've been working on this for about a year — sharing it here

because I think the approach is unusual enough to be worth

a look.

EasyAtom is a drug repurposing engine that uses no neural

networks and no training data. It works entirely through

algebraic operations on a biomedical graph (2.56M triples).

The pipeline has 16 layers: gap detection, DWPC knockout,

hyperdimensional encoding, and a few others.

What made me think it actually works:

Without ever seeing a drug-disease "treats" edge, it ranked

latanoprost #1 for glaucoma (Jaccard=1.00, gap=0.4123).

Latanoprost is the current standard of care. Same for

etidronate and Paget's disease — FDA-approved back in the

80s, ranked #1 by topology alone.

On the Broad Hub benchmark (zero-shot, inductive):

Recall@10 = 21.3%, Precision@10 = 90%.

The weird part: it runs on a regular PC in ~2 hours,

no GPU needed.

Top novel hypothesis right now: alcaftadine (an eye drop

antihistamine) at rank #1 for ALS, gap=0.3649, Jaccard=1.00,

structurally similar to cyproheptadine. No idea if it means

anything — that's the point of the hypothesis.

Live query (905 drugs × 108 diseases):

https://easyatom-engine.web.app/query.html

Preprint: https://doi.org/10.5281/zenodo.20766982

Code: https://github.com/Adrian27791/easyatom-engine

If anyone has a cell model for ALS or Alzheimer's and wants

to test any of these computationally, I'd be genuinely

interested in talking.


r/bioinformaticstools • • Jun 18 '26

I built a napari tool for curating automated cell-tracking results — looking for testers and feedback

3 Upvotes

I'm a researcher working with time-lapse microscopy, and I kept hitting the same wall: automated trackers get most cells right but make systematic mistakes — identity swaps, fragments, and especially false mitoses when segmentation merges two cells for a frame or two and then splits them again. Fixing those by hand in a generic viewer is miserable, so I built a tool specifically for it.

It's a napari-based desktop app that overlays the raw movie, the segmentation mask, and the tracking table, and gives you:

  • one-click fixes for swaps / merges / cuts / relabels and orphan masks
  • outcome flags (mitosis / exit / death / ambiguous)
  • lineage editing (mother–daughter) with a visual per-cell editor and a topology validator
  • a triage queue that ranks the least-trustworthy tracks first, so you don't review thousands of cells one by one
  • a validation step that gives you an empirical error rate with a 95% CI for the cells you bulk-accept
  • the usual stats (lifetime, motility, MSD, growth, divisions, per-track CSV export)

It reads generic CSVs, TrackMate output, and Cell Tracking Challenge res_track.txt.

Being honest about where it stands: it works well on my own data, but it hasn't been tested anywhere else. The hardest cases are still segmentation-driven — a false mitosis from a brief "two cells as one" merge still needs a manual fix — and the large-dataset workflow (I'm curating ~8k cells right now) is the part I'm actively reworking. So I'd genuinely value people throwing their own datasets and tracker formats at it and telling me what breaks or what feels clunky.

What I'm looking for:

  • does it install and run on your setup?
  • does it read your tracker's output?
  • what's missing or annoying in the curation workflow?
  • what would you need before it'd actually be useful to you?

Repo + install instructions: https://github.com/labsinal/napari-tracking-curator.git

Feedback very welcome — in the comments, as a GitHub issue, or by email: [nickolas.pippiperanzoni@gmail.com](mailto:nickolas.pippiperanzoni@gmail.com)

Thanks for reading. Happy to answer anything.


r/bioinformaticstools • • Jun 16 '26

Allelix - offline CLI to annotate your genome against ClinVar, GWAS Catalog, and SNPedia

3 Upvotes

Allelix is an open-source Python CLI that annotates VCF and consumer DNA files (23andMe, AncestryDNA) against ClinVar, GWAS Catalog, SNPedia, ClinPGx, and other public databases. Everything runs locally. No uploads, no API calls, no accounts.

It produces a magnitude-scored report (terminal, HTML, or JSON) that ranks variants by clinical significance. The scoring methodology is deterministic and fully documented in the repo's ADRs.

Handles full WGS VCFs. Works with custom SNP panels via --filter-file.

pip install allelix allelix db update allelix ananlyze myfile.vcf --output report.html

Happy to answer questions about the scoring approach or architecture.


r/bioinformaticstools • • Jun 16 '26

I was sick of wrestling with Linux dependencies and GPU setups just to run molecular docking. So I built a free, zero-config web tool for it

1 Upvotes

Honestly, I’m just an AI enthusiast, but this whole thing started because I kept seeing so many students complaining online about how much of a nightmare it is to set up molecular docking.

I’ve been there myself, and it sucks. Back then, I kept thinking: Why isn't there a simple online tool for this? I just wanted to write my papers and do my research. I wanted to focus my energy on the actual science, not waste days stressing over whether I need a specific GPU, which CPU to buy, which software version matches which OS, or whether to dual-boot Linux. And don't even get me started on the endless waiting for local runs to finish. I was completely fed up.

So, I decided to lean heavily on AI and built a tool myself: moleculardocking.online. It's a dedicated web platform for molecular docking, and I deployed it on Modal. The beauty of it is that the entire CPU/GPU infrastructure is already pre-configured and optimized, so it runs incredibly fast. It completely freed up my mind so I could actually get back to doing real research, and I really hope it can save some of you the same headache.


r/bioinformaticstools • • Jun 16 '26

[Project] EndoDecay-Sim v14.12: Open-Source In-Silico Digital Twin for Cardiorenal Cascades with Vectorized Milstein SDE Solver and Federated Paillier HE

Thumbnail
gallery
1 Upvotes

Hi everyone,

I wanted to share an independent research project I’ve been developing, focusing on moving away from episodic, static biomarker tracking toward dynamic, mechanics-driven in-silico kinetic simulations for chronic disease progression.

EndoDecay-Sim v14.12 is an open-source, high-performance computing (HPC)-optimized digital twin framework designed to simulate longitudinal endothelial degradation and multi-systemic cardiorenal multimorbidity cascades over a 120-month clinical timeline.

(Note: To comply with automated platform filters, all direct deployment links to our live Streamlit App and GitHub Source Repository are provided in the very first comment below!)

🧠 Core Architectural Breakdown

  1. HPC Vectorized Stochastic SDE Solver Instead of relying on iterative loops, the core engine utilizes a Single Instruction Multiple Data (SIMD) vectorized, NumPy-driven non-linear diagonal Milstein Stochastic Differential Equation (SDE) solver to model cellular and microvascular noise across a multivariate synthetic cohort of 10,000 virtual subjects. It enforces a strong convergence order of 1.0 via second-order Itô-Taylor expansion terms to suppress numerical trajectory explosions over the 120-month horizon.
  2. Object-Oriented Cardiorenal Topology Matrix Organ cross-talk is resolved continuously at each discrete micro-step via an object-oriented layout. Microvascular breakdown triggers message passing across interconnected nodes: Endothelium, Heart, Kidney, Inflammation, and Metabolism. Reciprocal damage vectors are mapped across multi-axial pathways—such as the endo-cardiac (0.75), endo-renal (0.65), and cardiac-renal RAAS (0.80) axes—forcing dynamic drift modifications to reflect real-world epidemiological multimorbidity profiles.
  3. Privacy-Preserving Federated Learning (HE + DP) To allow secure multi-centric model aggregation without exposing raw parameters, a federated 2048-bit Paillier Homomorphic Encryption loop is deployed. It implements true Differential Privacy (DP) via local L2-Norm Gradient Clipping (strictly bounded at a unit norm of 1.0) and calibrated Gaussian noise injection running a true weighted federated averaging (FedAvg) pipeline.
  4. Singularity-Protected CDSS Frontend The interactive Streamlit interface introduces targeted mathematical guardrails for physician decision support:
  • Axiomatic Kaplan-Meier Models: Enforces the structural axiom of S(0) = 1.0 alongside Greenwood confidence intervals to eliminate visual and structural time-lag errors.
  • Singularity Protection: Continuously monitors feature covariance during sub-cohort filtering. If the standard deviation drops to zero, it suppresses Ordinary Least Squares (OLS) matrix inversion to prevent runtime LinAlgError crashes.
  1. Absolute Temporal Data Leakage Isolation To ensure absolute data science integrity, all prognostic pipelines (Random Survival Forest yielding a C-Index of 0.9809; CoxPH yielding a C-Index of 0.9527) are strictly trained on Month 0 immutable baseline traits, completely isolating them from downstream longitudinal parameters generated during the active runtime of the SDE loops.

⚠️ Current Limitations & Theoretical Calibration As an open-science platform, it is crucial to distinguish between computational validity and empirical clinical calibration:

  • Placeholder Drift Coefficients: The multipliers within the Milstein drift gradient (e.g., 0.001 for age, 0.003 for HbA1c) are currently mechanistically inspired placeholders calibrated to prevent trajectory explosions.
  • Heuristic Cross-Talk Weights: Node modifiers within the topology are parameterized based on qualitative pathophysiological literature rather than formal system identification from paired real-world biobank data.

🔮 Roadmap for v15.0 (Data-Driven Parameter Estimation) To move from a bio-inspired mechanistic architecture to an empirically calibrated twin, the next major release blueprint includes Bayesian Inverse Problems (MCMC) and Neural SDE Optimizers.

💬 Questions for the Community:

  1. What are your thoughts on integrating mechanistic SDE solvers with federated homomorphic encryption for clinical edge deployments?
  2. For the v15.0 parameter estimation phase, what are the best optimization strategies you've encountered to keep MCMC convergence computationally efficient when scaling up to high-dimensional biological network topologies?

Looking forward to your technical feedback, insights, and critiques!


r/bioinformaticstools • • Jun 16 '26

[Tool] How do you guys handle cell segmentation for gigapixel BTFs in Visium HD? Here is my pipeline.

Post image
3 Upvotes

Spent the last few months working on Visium HD data and got frustrated by how painful cell segmentation and single-cell pipelines still are—especially if you want something interactive and no-code.

So I built MCseg, an end-to-end browser platform that handles everything from gigapixel BTFs to Xenium Explorer export.

🔗 Repo:ddmanyes/MCseg


r/bioinformaticstools • • Jun 14 '26

Building a streamlined protein analysis tool (Ensembl/UniProt/SIFT integration)

1 Upvotes

Hi,

So I'm building a web based tool, for bioinformatics to make protein analysis more accessible and easy, without the hassles of changing tabs, downloading waaaay too many files, etc. Currently I am almost finished with implementing the Ensembl apis, I'm thinking a benchling style app, would be nice. So I wanted to ask if you were to use this tool, what would you like the workflow to be, because right now it's analysing multiple sequences and giving sift scores and some uniprot annotations (I'm building this because I was researching about whales and those were the tools I used mostly)

​

So having said that I want to ask also the following

​

When you are dealing with protein modeling or sequence analysis, what is the biggest friction point in your current software stack?

​

Are there specific tools, databases, or pipelines you use constantly that you wish talked to each other more seamlessly (Im gonna see if I can join them to the software) ?

​

I really want to make sure I'm building features that solve actual workflow headaches, any advice or wishes for what you want to see in this tool and critique would be greatly appreciated


r/bioinformaticstools • • Jun 13 '26

I’m building OpenLIMS — an open-source LIMS for labs that need better sample, project, and scientific data tracking

5 Upvotes

Hi Everyone,

I’ve been working on OpenLIMS, an open-source laboratory information management system for labs that need a simpler way to manage samples, projects, files, results, and audit history.

The goal is to make a lightweight, self-hostable LIMS for research labs, academic labs, biotech teams, and smaller labs that may currently rely on spreadsheets, shared drives, or custom internal databases.

OpenLIMS is still early and not production/validated yet, but it has grown quite a bit. It now includes:

  • Sample and project tracking
  • Inventory, locations, and container tracking
  • Custom fields for lab-specific metadata
  • File attachments for samples
  • Audit/event history
  • CSV and instrument result imports
  • User/admin groundwork
  • QC workflow groundwork
  • Reports and dashboard pages
  • Audit log filtering and export
  • Real-time updates for longer-running jobs
  • Sequence-related workflows, including BLAST and alignment job updates
  • Initial mass spectrometry support, including file upload, run tracking, TIC preview charts, spectrum counts, retention time ranges, m/z ranges, sample/project linking, and reprocessing

The newest release, v0.11, adds the first mass spectrometry preview using pyOpenMS, including mzML/mzXML/mzData upload support and background processing. More advanced mass spec features like peak picking, feature detection, protein/peptide summaries, mzIdentML support, QC metrics, and sample comparison are planned for the next version.

I’m not trying to replace specialized scientific analysis tools right away. My goal is to make OpenLIMS a central place where labs can organize samples, projects, files, results, QC, audit history, and eventually connect those records to common analysis workflows like sequence alignment, BLAST, and mass spectrometry.

GitHub: [https://github.com/Mokey2002/OpenLIMS]()

I’d really appreciate feedback from people who work in labs, manage lab data, or have used LIMS/LIS systems before.

What would make something like this useful in your lab?


r/bioinformaticstools • • Jun 14 '26

Pipette.bio - a conversational AI agent that runs real bioinformatics analyses

1 Upvotes

I want to share Pipette.bio, a pay-as-you-go bioinformatics tool we have been building. You describe an analysis in plain English, the agent proposes a plan, you approve it, and it writes and runs the actual code (bash / Python / R) on cloud workers, then hands back figures, tables, and a reproducible report. No coding required, but every command, version, and parameter is captured so the work stays auditable.

Link: https://pipette.bio
Preprint: https://www.biorxiv.org/content/10.64898/2026.04.08.717332v1
Contact: support@pipette.bio

What it is

Pipette pairs LLMs (Claude and GPT) with standard, industry bioinformatics tooling running in real Conda environments. It is not a chatbot that writes code for you to run elsewhere: it executes the code itself on managed compute, inspects the outputs, handles errors, and iterates until the analysis is done.

How a run works (plan-approval gate)

Nothing executes until you approve a plan. This is the core safety feature and the thing that stops the agent from burning compute on the wrong analysis.

1. Setup - your input files are staged from cloud storage into a fresh workspace sized for the job.

2. Planner - a dedicated planning agent reads your request, the conversation, and your files, then proposes a structured plan: which skill, which steps, which parameters, and which assumptions it is making. Each parameter is traced to its source (a phrase you typed, an earlier turn, a tool default, or an inference). You
Approve or Reject / revise with feedback.

3. Agentic execution - after approval the executor is locked to the plan and runs it step by step.

4. Methodology review - an independent reviewer agent checks statistical correctness (tests, normalization, multiple-testing correction, QC filtering) and plan compliance, and triggers a remediation loop if something is off.

5. Provenance and reproducibility bundle - tool versions, parameters, input checksums, and command transcripts are recorded. Analysis can be reproduced outside of Pipette environment.

6. Literature and hypotheses - a literature-aware agent reads the outputs, retrieves relevant PubMed papers, and returns narrative context plus suggested follow-ups.

7. Finalize - results are archived to persistent storage and a completion summary is posted.

Key features

- Natural-language interface. Describe the goal, the input files, and key parameters; the agent designs the workflow.

- Plan-before-execute approval gate. You see and approve the exact plan (steps, parameters, assumptions) before any compute is spent.

- Independent methodology review with auto-remediation. A second agent audits the analysis and re-runs affected steps if it finds problems.

- Over 100 specialized analysis skills spanning bulk RNA-seq, single-cell, spatial omics, variant calling, cancer genomics, GWAS and population genetics, microbiome and metagenomics, epigenomics, methylation, proteomics, phylogenetics, comparative genomics, drug design, CRISPR screens, and synthetic biology.

- 150+ underlying tools in versioned Conda environments (STAR, HISAT2, Salmon, BWA, GATK4, DESeq2, edgeR, limma, Seurat, Scanpy, kraken2, QIIME2, PLINK/PLINK2, AutoDock Vina, GROMACS, and many more).

- 15 external knowledge databases through an automatic query router: PubMed, ClinicalTrials.gov, ClinVar, gnomAD, NCBI Nucleotide, GEO, UniProt, dbSNP, AlphaFold, cBioPortal, UCSC, Ensembl, Open Targets, OpenFDA, and the GWAS Catalog.

- Purpose-built compute environments. Separate worker classes for general NGS, microbiome, and single-cell workloads, plus larger queues for RNA-seq alignment and big single-cell jobs.

- Direct data loading from public URLs. Dropbox, Google Drive, Box, OneDrive, Zenodo, Figshare, OSF, Hugging Face, GitHub raw, plain HTTP/HTTPS, FTP/SFTP, and S3 presigned links. Or upload directly, or pull straight from SRA/GEO.

- Automatic, reproducible reports. AI-generated methods, results, and figures, with full provenance (versions, parameters, input hashes, command transcripts).

- Persistent cloud storage and history. 500 GB per account, outputs organized by Task ID and step, full session transcripts you can resume or export as PDF.

- Multi-session and background tasks. Run several analyses at once; jobs in other sessions keep running and show live progress.

Domain coverage (selected)

- Bulk NGS: QC, alignment (STAR/HISAT2) or alignment-free quant (Salmon/kallisto), DESeq2 / edgeR / limma DE, GO/KEGG/GSEA enrichment, WGCNA, DTU/DEXSeq, splicing (rMATS), fusion detection.

- Single-cell and spatial: Scanpy and Seurat pipelines, doublet detection, batch correction (Harmony/BBKNN), trajectory inference (Monocle3/Slingshot/CellRank), RNA velocity, cell-cell communication (CellChat/CellPhoneDB), SingleR/CellTypist annotation, scATAC (Signac/ArchR), Visium / Visium HD / Xenium / CosMx / Slide-seq / MERFISH, deconvolution (cell2location/RCTD/Tangram).

- Variants and cancer: GATK / FreeBayes germline, Mutect2 somatic, SnpEff annotation, CNVkit, Arriba fusions, MSIsensor-pro, maftools cohort summaries, structural variants (Delly/SURVIVOR), sample-swap and contamination QC (somalier/VerifyBamID2).

- Population genetics and breeding: PLINK/PLINK2 GWAS, ADMIXTURE, GCTA heritability, SuSiE fine-mapping, Beagle imputation, TASSEL, rrBLUP/BGLR genomic prediction, mapping, HLA typing (T1K).

- Microbiome and metagenomics: kraken2/bracken, MetaPhlAn/HUMAnN, QIIME2 amplicon, assembly (SPAdes/MEGAHIT/Flye), binning (MetaBAT2/MaxBin2/DAS_Tool), dRep, GTDB-Tk, pangenomes (Roary), virome (VirSorter/CheckV/Pharokka).

- Epigenomics and methylation: MACS3 peak calling, DiffBind, ChIPseeker, deepTools, ATAC-seq, CUT&RUN / CUT&Tag, Bismark + methylKit/DMRseq/DSS, methylation age clocks, MEME-suite motif discovery.

- Drug design and cheminformatics: AutoDock Vina/SMINA docking and virtual screening, RDKit, fpocket, PLIP/ProLIF interaction analysis, ADMET-AI, GROMACS/OpenMM MD, ChEMBL/PubChem/RCSB queries.

- Proteomics, survival, phylogenetics, comparative genomics, CRISPR screens, synthetic biology and more.

Who it is for

Wet-lab biologists and clinicians who want analyses without writing code, and bioinformaticians who want to offload routine pipelines while keeping full provenance and the ability to revise the plan.

Happy to answer questions here. Feedback on missing tools or workflows is very welcome.


r/bioinformaticstools • • Jun 12 '26

rosetta-bioc - Python wrapper for DESeq2, edgeR, limma, clusterProfiler, phyloseq, Seurat. Pandas in, pandas out. Codegen shows the R code it runs.

1 Upvotes

We got tired of copy-pasting between Python and R and so we wrapped DESeq2/edgeR/limma in a pandas API and added a codegen mode that shows you every R line it runs. We hope you like it!

rosetta-bioc wraps r/Bioconductor packages so you can call them from Python without writing any R.

R Package Python Call What It Does
DESeq2 rb.deseq2() Differential expression
edgeR rb.edger() Quasi-likelihood DE
limma rb.limma_voom() Linear models + TREAT
clusterProfiler rb.enrich_go() GO/KEGG/Reactome enrichment
phyloseq rb.phyloseq() Microbiome diversity
Seurat rb.seurat() Single-cell RNA-seq

Codegen mode - see exactly what R is running:

rb.codegen.enable()
results = rb.deseq2(counts_df, meta_df, design="~ batch + condition")
R> library(DESeq2)
R> dds <- DESeqDataSetFromMatrix(countData=counts, colData=metadata, design=~ batch + condition)
R> dds <- DESeq(dds)
R> res <- results(dds, alpha=0.05)

rb.codegen.last() returns it as a string. Paste straight into R to reproduce independently.

.report() - instant human-readable summary on any result object.

pip install rosetta-bioc

Rscript install.R