Skip to content
Official MCP RegistryListed

com.seqbench/workbench

Hosted DNA/RNA/protein tools: primers, oligos, PCR, cloning, CRISPR, alignment, batch & pipelines.

First seen 2 Oct 2026. Evidence as of 2 Oct 2026.

144
Tools
From an anonymous probe
1
Source listings
Each with its own history
0
Recorded changes
Since first seen

Tools

ToolDescriptionBehaviour
alphafold_lookupLook up a UniProt accession in the AlphaFold Protein Structure Database (CC-BY 4.0). Returns confidence, model version and structure file URLs, or {found:false} when no prediction exists for that accession.Read-only
annealing_temperatureAnnealing temperature and a full cycling program for a primer PAIR, under the rule the polymerase's own vendor publishes — which is not one rule, and is only defined against the vendor's own Tm. For NEB products the Tm is computed exactly as NEB's Tm calculator computes it (SantaLucia 1998 nearest-neighbor with NEB's salt correction, at that product's own buffer and primer concentration — the same primer's Tm moves ~10 °C between NEB buffers) and NEB's calculator rule is applied to it: Q5 at Tm + 1, Phusion at 0.93 × Tm + 7.5 (both ABOVE the Tm), Taq/OneTaq/LongAmp at Tm − 5, Vent/Deep Vent at Tm − 2 above 20 nt, with NEB's ceilings. Toyobo KOD One is Tm − 5 at the salt and oligo concentration its manual names, Takara PrimeSTAR Max a FIXED 55 °C the primer Tm plays no part in (its own 2(A+T)+4(G+C)−5 sets only the annealing time), and a generic Taq/Pfu row applies the Tm − 3 to − 5 rule of thumb to a default-condition nearest-neighbor Tm. Returns denaturation/annealing/extension steps, cycle count, magnesium guidance, the Tm basis and vendor document behind each figure, and the same primers under every other polymerase in the table for comparison.Read-only
aso_designDesign antisense-oligonucleotide (ASO) gapmers against an mRNA target: scans candidate sites, builds the antisense oligo in the standard 5-10-5 architecture (chemically-modified wings, central DNA gap for RNase H1, phosphorothioate backbone), and screens each for known liabilities (G-quadruplex motifs, CpG immunostimulation, self-complementarity, GC extremes). No transcriptome-wide off-target search.Read-only
assembly_outcomesEnumerate the specific wrong plasmids a multi-part Golden Gate or Gibson assembly can produce — a part dropped, inverted, duplicated, two parts swapped, the backbone self-circularized — as full sequences, ranked by how few independent mis-ligations each needs. Golden Gate outcomes are annotated with the MEASURED overhang cross-talk they would have to exploit (Potapov/Pryor ligation data). Feed the result to diagnostic_digest to pick a screening enzyme. Reports no probability per outcome: the ligation data does not measure transformation or vector background.Read-only
band_tracebackExplain a band you measured on a gel. Given the template, both primers and the observed size, it enumerates every pair of priming sites — including a single primer priming both strands — that would give a product that size, and ranks them by how much of each primer's 3' end matches without interruption, which is what decides whether a mispriming event can extend at all. Reports no yield and assigns no share of the band: the band is the input, not the output. Says plainly when nothing on this template explains the size, and what that points to instead.Read-only
barcode_auditTake a set of barcodes and report the minimum edit distance in it, which pairs sit at that distance, any duplicates, the GC range and the longest homopolymer. The pairs are the point: 'minimum distance 2' is a number, and 'these two barcodes are one substitution apart' is something to fix. Use before committing a pool you inherited or assembled by hand — a single close pair makes two constructs indistinguishable in the data and nothing downstream can detect it.Read-only
barcode_designBuild a set of DNA barcodes that are pairwise at least a chosen edit distance apart, so a read error cannot turn one barcode in the pool into another. Honours a GC window, a homopolymer cap and excluded motifs (the restriction sites you clone with) during the search rather than filtering afterwards, and can extend a set you already have rather than replacing it. Reports the distance ACHIEVED — re-derived over the finished pool, not assumed from the search — plus how many read errors that distance lets you correct and detect. Deterministic given a seed. Returns fewer barcodes than asked, with a reason, when the constraints leave no room.Read-only
base_edit_quantQuantify CBE/ABE base editing from a pair of Sanger traces — an unedited control and the edited pool — without NGS. At each editable position in the activity window the edited trace is treated as a mixture of the unedited and converted peaks, and the control's OWN alt-channel signal at that same position is subtracted as background, because dye crosstalk is position- and context-dependent and a global constant would be wrong per position. Significance comes from a null built from the same sample (every control position outside the window carrying the same base), so the threshold adapts to the run's chemistry instead of being hardcoded. Returns per-position percentages with z-scores, the target and its bystanders, the background distribution (including a robust estimate of its spread and a count of its outliers), and the noise floor — the percentage the background alone reaches, or null when the run's own null has no spread to derive one from. A window position whose control already carries the converted base is reported but not quantified, because the (1 − b) rescale amplifies error by 1/(1 − b) and turns a 0.1-point wobble into half the pool. Locate the window with an editor id plus the protospacer, or give it explicitly. Blind to indels, which shift the trace rather than mixing a base. PREDICTED, NOT MEASURED. None is published for this implementation. Every run instead reports what it rests on: the background mean, sd, robust (MAD-based) sd, outlier count and n, and a noiseFloorPercent that says how much apparent editing the background alone reaches at the significance threshold — null, rather than 0, when the run has no null with any spread to derive a limit from. This implementation's recovery of known synthetic mixtures (to under a percentage point) is deliberately NOT offered as validation — it tests the arithmetic and the coordinate handling, not whether the linear mixture model fits a real capillary trace. Two properties ARE characterized. The bias is directional and one-sided: the (1 − b) rescale assumes a fully converted position would read as alt fraction 1.0, which real chemistry does not reach, so percentages run low by roughly the crosstalk fraction (order 5-10% relative at typical 3% bleed). The variance is not constant across positions: dividing by (1 − b) amplifies the error in a and b by 1/(1 − b), so the variance of the estimate goes as 1/(1 − b)². Positions whose control alt fraction b exceeds 0.25 are therefore reported but NOT quantified, which bounds that amplification at 1.33x on everything the tool does quantify. Valid for: A pool edited by a cytosine or adenine base editor, read on the same amplicon and chemistry as an unedited control that starts within ±40 bases of it, where the editing is a SUBSTITUTION. Not valid for indels — a base editor also makes them, and an indel-bearing allele shifts the downstream trace so that it degrades the fit at every window position rather than showing up anywhere. Not valid where the control read already carries the alt base at a window position (a pre-existing variant, or the wrong control), since then there is nothing left to correct against — that condition is DETECTED rather than only described: a window position whose control alt fraction exceeds 0.25 is reported with an excludedReason and an editedPercent of 0 instead of a number, warned about, and failed by the hard control-supports-the-window-base-calls gate check. Reported percentages are pool-level: Sanger sees the superposition, so which allele carries which combination of bystander edits is not recoverable from it at all.Read-only
base_editing_designDesign cytosine (CBE, C→T) or adenine (ABE, A→G) base-editing gRNAs for an SpCas9 target: for each NGG gRNA it reports every editable base inside the editor's activity window, flags bystander edits (more than one editable base in the window), and — with a CDS reading frame — classifies each edit's amino-acid consequence (silent / missense / nonsense / stop-loss). Bystander-free guides are ranked first. Handles both strands (a C→T on the protospacer of a reverse-strand guide is reported as the forward-strand G→A).Read-only
batchRun one batchable SeqBench tool over many records. Returns a table of per-record results and typed failures.Read-only
blast_pollCheck a BLAST search submitted with blast_submit. Returns {ready:false} while it is still running; once ready, returns each hit's accession, title, organism, best HSP (bit score, expect, percent identity, span on both query and subject) and the percentage of the QUERY it covers, with overlapping HSPs merged so coverage cannot exceed 100%. A request id NCBI no longer recognises is reported as either a typo or an expired search, which is the same answer it gives for both.Read-only
blast_submitSubmit a sequence to NCBI BLAST and get a request id back immediately — poll it with blast_poll. The program is inferred from the sequence (nucleotide to blastn, protein to blastp) unless you name one, and a query whose type does not match the named program is refused rather than silently returning nothing. Databases: blastn: nt, refseq_rna (default nt); blastx: nr (default nr); blastp: nr, swissprot (default nr); tblastn: nt (default nt). Up to 20,000 residues. Searches typically take 20-60 seconds; NCBI asks that a single search is polled no more than once a minute.Read-only
centrifuge_conversionConvert between rotor speed (RPM) and relative centrifugal force (×g) for a given rotor radius, in either direction. The conversion is exact physics — RCF = ω²r/g — and the only judgement in it is which radius: a protocol's ×g conventionally means the rotor's MAXIMUM radius, and using the wrong one scales the force by the ratio of the two.Read-only
characterize_sequenceOne-paste 'tell me everything': auto-detects DNA/RNA/protein, then reports composition, ORFs, single-cutter enzymes, end primers or protein properties, plus a BLAST link.Read-only
cloning_diagnoseWork out why a cloning experiment failed: no colonies, every clone empty vector, or no PCR band. Takes your design (method, parts, enzymes, primers, host methylation state) plus what you actually observed (colony counts on the plate and on each control, screening tally, band sizes, whether the ladder ran) and returns causes ranked by evidence — each with the deterministic fact from the design or the observation that implicates it, the cheapest observation that would separate it from the next candidate, and the next experiment. Causes the observations eliminate are reported as eliminated, naming the observation that did it; causes the design makes impossible are not listed. No probability is computed anywhere — the ordering is of evidence, not of likelihood, and `ranking.evidenceBased` says so when the inputs separate nothing.Read-only
cloning_next_observationRank every observation you have not yet made by how many open causes it settles WHICHEVER WAY IT COMES OUT, then return the smallest set of observations that settles all of them. Takes the same arguments as cloning_diagnose. The ranking is computed by re-running the diagnosis at each possible outcome of each observation and intersecting, so every count is a worst case rather than an average — and causes that no observation can settle are named, because those need a different experiment rather than more observing. No probabilities anywhere.Read-only
cloning_simulateAssemble fragments by Gibson/overlap, Golden Gate (Type IIS), restriction–ligation (sticky or blunt), TOPO/TA, LIC or SLIC (T4-polymerase chew-back) or In-Fusion/CPEC, returning the product, the junctions and — for the primer-design methods — the junction primers. Each method is modeled as its own chemistry rather than as one product model with different labels: LIC's chew-back stops at the first occurrence of the single dNTP supplied, so a tail carrying that base stops it early and a tail without one lets it run past the junction, and both are refused with the offending base and position named.Read-only
codon_adaptation_indexCodon Adaptation Index (CAI) and per-codon relative adaptiveness of a CDS against an expression host, with rare-codon and GC3 analysis.Read-only
codon_optimizeCodon-optimize a protein (or coding DNA) for an expression host by picking the most-frequent codon per residue.Read-only
construct_autofixIteratively substitutes synonymous codons to resolve unwanted restriction sites (domestication for Golden Gate), homopolymers, tandem repeats, predicted secondary structure, cryptic RBS/polyA motifs and hidden alternate-frame ORFs that construct_qc flags — without changing the encoded protein (verified). Does NOT touch premature stops or GC extremes; re-run construct_qc afterward to confirm. A native TypeScript alternative to a constraint-solver sidecar.Read-only
construct_qcLint a coding DNA sequence for premature stops, internal RBS/polyA motifs, unwanted restriction sites, GC extremes and repeats.Read-only
cpg_islandsFind CpG islands by either published definition: Gardiner-Garden & Frommer 1987 (GC > 50%, observed/expected CpG > 0.6, >= 200 bp — the original, which also calls many Alu repeats) or Takai & Jones 2002 (GC >= 55%, obs/exp >= 0.65, >= 500 bp — written to exclude them). Windows are merged only where the merged region STILL meets the criteria, so every island returned satisfies the definition it is reported under. Reports the CpG DINUCLEOTIDE count separately from GC content, because conflating the two is the usual way this number goes wrong.Read-only
crispr_grna_designFind and score candidate guide RNAs (protospacer + PAM) in a target DNA for common nucleases (SpCas9, SpCas9-NG, SaCas9, Cas12a). PREDICTED, NOT MEASURED. No held-out skill statistic is claimed. Both are pre-2016 models, superseded by Rule Set 3 (DeWeirdt et al., Nat Commun 2022) and by DeepHF. Rule Set 3 IS now shipped here, as crispr_ontarget, and unlike these two it carries a held-out calibration: on an independent tiling library, 87.7% / 74.9% / 82.0% of its lowest-scoring guides landed in the bottom two activity quintiles. Prefer it for SpCas9 with an NGG PAM; DeepHF is still not shipped. Treat these two as a ranking aid, not an efficiency prediction. Valid for: SpCas9 with an NGG PAM and a 20 nt spacer, and only when enough genomic flanking context is present to build the model's 30-mer / 35-mer window — both scores are null rather than padded otherwise. Nothing is predicted for SaCas9, Cas12a or SpCas9-NG.Read-only
crispr_hdr_donorBuild an HDR donor (homology arms flanking an edit) from a target sequence and either an explicit edit window (editStart/editEnd) or a guide's cut site (guideStart/guideEnd/guideStrand/nuclease — SpCas9-family only; Cas12a's staggered cut needs an explicit editStart/editEnd). Also designs genotyping primers spanning the edit site on the original sequence (a real size-shift or sequencing target to confirm the edit), reusing the same primer-design engine as primer_design.Read-only
crispr_offtarget_checkScreen a guide's protospacer for off-target sites (protospacer match + valid PAM, both strands) against a small curated set of common lab reference genomes (see genomesChecked) — NOT a whole human/mouse genome search. For SpCas9 with a 20 nt spacer each site also gets a Doench 2016 CFD score, so sites are ranked by predicted cut likelihood rather than by mismatch count alone, and the guide gets an aggregate specificity. Use this the same way primer_specificity is used: a useful sanity check within the covered organisms, not a clearance guarantee for a mammalian expression host. PREDICTED, NOT MEASURED. Best of the common off-target scores on the authors' GUIDE-Seq comparison, at Pearson r = 0.40 over 9 guides and 402 sites (vs CCTop 0.31, Hsu-Zhang 0.26) — a useful ranking, not a reliable magnitude. Weights were measured for SINGLE mismatches; multiple mismatches are multiplied, and Listgarten et al. 2018 note the training data never contained a mismatch and an alternative PAM together, so that combination is extrapolation. Valid for: SpCas9 with a 20 nt spacer, which is the only case scored — every other nuclease returns null rather than a number from the wrong enzyme's table. One locus, one cell line. Substitutions only: DNA/RNA bulges are neither searched nor scorable, and a low score is not a claim that a site is safe.Read-only
crispr_ontargetScore SpCas9 guides for on-target activity with Rule Set 3 (DeWeirdt et al., Nat Commun 2022) — the current successor to the Doench 2014 and CRISPRscan scores crispr_grna_design reports. Give a target sequence to find and score every NGG guide in it — on both strands unless `searchReverseStrand` says otherwise — or give 30-mer contexts directly. The tracrRNA MATTERS and is not cosmetic: Rule Set 3 models it as a feature, and the paper measures its accuracy dropping when the wrong one is specified. Returns a z-scored activity, which ranks guides against each other; it is not a percentage and not a probability of editing. PREDICTED, NOT MEASURED. Held out six datasets (23,629 context sequences) from training; Rule Set 3 (Sequence) had the highest Spearman correlation on three of the six. On a separate tiling library generated for the paper, with every spacer any model had seen removed, it significantly outperformed all other models (p < 0.002) when the correct tracrRNA was given. Calibration on that independent set: of the lowest-scoring guides, 87.7% / 74.9% / 82.0% landed in the bottom two activity quintiles (Hsu / Chen / DeWeirdt tracrRNA), and of the highest-scoring, 77.0% / 69.0% / 75.5% landed in the top two. No single held-out Spearman is quoted here because the paper reports it per gene as a distribution rather than as one figure, and inventing a headline number from a figure would be the fit-residual mistake in a different costume. Valid for: SpCas9 with an NGG PAM and a 20 nt spacer, scored over a 30-mer (4 nt upstream + spacer + PAM + 3 nt downstream) — nothing is padded, so a guide without that flank is skipped rather than estimated. Knockout activity in mammalian pooled screens with a U6/Pol III promoter; the authors expect it to generalize less well to Pol II-transcribed sgRNAs, and in vitro transcribed sgRNAs (as used in zebrafish) are known to be poorly predicted by models trained this way. The tracrRNA must be the one you will actually use. This is Rule Set 3 (Sequence) only: the Sequence+Target model, which adds 6.7% to the Spearman correlation on average, needs per-gene conservation and protein-domain lookups and is not implemented here.Read-only
cross_dimerScreen two oligos for the most stable heterodimer (cross-dimer) between them.Read-only
diagnostic_digestPick the restriction digest that tells your intended construct apart from the wrong ones on a screening gel. Digests every candidate, works out which bands would actually resolve at the chosen agarose percentage (size ratio, the gel's resolving window, and whether a band is too faint to score), and ranks single enzymes — then buffer-checked pairs if no single one works. The criterion is separating the INTENDED construct from every alternative; telling the alternatives apart from each other is reported as a bonus. Get the alternatives from assembly_outcomes.Read-only
dna_molarityNucleic-acid quantity conversions: molar mass, amount (pmol/nmol), molar and mass concentration, and copy number, from mass ± volume and either a length or a sequence.Read-only
dot_plotWord-match dot plot between two sequences, on both strands, returned as diagonal RUNS rather than points. This is the view that shows what a single best alignment hides: internal repeats, tandem duplications, and inversions — an inversion appears as a run on the falling diagonal and as nothing at all in a normal pairwise alignment. Compare a sequence with itself to map its own repeats. Word size is chosen from the input lengths when you do not set one, because too short a word fills the plot with chance matches that the eye reads as homology.Read-only
double_digestRecommend a single NEB buffer (and flag caveats) for digesting with two enzymes in one tube.Read-only
editing_plate_quantifyQuantify a whole plate of edited samples against ONE untreated control trace and return a single sortable table — the plate-scale form of sanger_indel_spectrum, base_edit_quant and sanger_knockin_quant, chosen with `mode`. Each sample gives one row keyed by its id, carrying the headline number for that mode (edited fraction / editing at the target base / intended knock-in percentage), the fit-quality numbers behind it (R², or the background n and noise floor for base mode), and fitAdequate — the single-sample tool's own gate verdict on that row, so the plate cannot drift from the per-well answer. Failure is isolated per well: a sample whose read is short, mismatched or unfittable becomes a failed ROW with its error message and the other 95 still come back, while an error about the control trace, the mode or the work ceilings throws, because it is wrong for every row. Duplicate sample ids are suffixed (against the whole plate, so the suffix never lands on another well's name) rather than merged. Arguments are strict: an argument belonging to another mode, an unknown argument, an out-of-range limit, and an `offset` override (which is a property of one pair of reads, not of a plate) are all rejected rather than ignored or clamped, because at plate scale a substituted setting rewrites every row identically and nothing in the table looks odd. Returns the rows in input order, a tally, and a CSV. Comparing two wells' percentages is only meaningful when both rows are fitAdequate, which is why the plate summary is computed over those rows alone. PREDICTED, NOT MEASURED. None is published for this implementation, and being a batch does not soften that: each row is exactly the claim the corresponding single-sample tool makes. What every row instead reports is what it rests on — R² for the indel and knock-in modes, the background n, sd and noise floor for base mode — plus fitAdequate, which is the single-sample tool's OWN gate verdict on that row rather than a threshold re-invented here. Recovery of known synthetic mixtures is deliberately NOT offered as validation: it tests the arithmetic and the plate plumbing, not whether the model fits a real capillary trace, and for a knock-in with novel inserted bases it is circular, because a synthetic trace is built from the same idealised peaks the basis assumes. Quoting it would be the mistake rbs_predict made when it shipped a calibration residual as held-out skill. Valid for: One control read and a set of edited reads that are all the SAME amplicon, chemistry and primer as that control, with each read extending well past the edit site. COMPARING TWO WELLS' PERCENTAGES IS ONLY MEANINGFUL WHEN BOTH ROWS ARE fitAdequate: a percentage from a badly fitting well is not a smaller number than one from a well that fitted, it is a different kind of statement, and the plate summary here is therefore computed over the adequate rows only. Ranking wells also assumes they differ only in the variable under test — the same control is subtracted from all of them, so a well whose read started 30 bases later or whose reaction was dirty carries that difference into its number. Mode-specific limits carry over unchanged: indel mode is blind to substitutions, base mode is blind to indels and its percentages run low by roughly the crosstalk fraction, and knock-in mode cannot separate an intended pure deletion from an NHEJ deletion of the same length at the same site.Read-only
elsa_capacityReport how many promoters, sgRNA handles and neutral spacers are available at each maximum-shared-length threshold, and therefore the longest extra-long sgRNA array that can be built without reusing a part. Ask this before designing: the answer is a property of the measured parts collection, not of your guides, and it is the constraint that decides the design. Deterministic — it is a selection over a fixed table.Read-only
elsa_designBuild a multiplexed CRISPR array that expresses many sgRNAs from one cassette and shares no long stretch with itself. Twenty sgRNAs built the obvious way carry twenty copies of the same Cas9 scaffold, promoter and terminator — a construct that recombines in the cell and that synthesis vendors refuse — so this draws each transcription unit's parts from the measured non-repetitive collection of Reis et al. (2019), selecting ACROSS promoters, handles and spacers at once rather than within each type, then re-measures the assembled molecule for repeats the junctions created. Returns the sequence, an annotated GenBank file and the pool usage. Takes guides; it does not choose or score them — use crispr_grna_design and crispr_offtarget_check for that.Read-only
export_echo_picklistGenerate a downloadable Beckman/Labcyte Echo acoustic-liquid-handler picklist CSV (columns: Source Plate Name, Source Plate Type, Source Well, Destination Plate Name, Destination Well, Transfer Volume, Name — the header row reproduced from PyEcho, a real open-source Echo-picklist generator) for the given PCR reactions, at the same well positions export_plate_layout assigns. Assumes a 5 uL Echo-scale PCR reaction (master mix 2500 nL, each primer 250 nL, template 250 nL, water 1750 nL) — a commonly used acoustic-dispensing miniaturization scale, not a universal standard; rescale the volumes for your own protocol. Source/Destination Plate Type uses a placeholder Echo plate-type code (384PP_AQ_BP) — replace with the exact type from your own Echo Plate Type Library. Each distinct template label gets its own well on the TemplateSource plate, row-major (A1, A2, … A24, then B1, …) across that 384-well source plate.Read-only
export_opentrons_protocolGenerate a downloadable Opentrons Python Protocol API (v2, OT-2) script that sets up the given PCR reactions on a 96-well PCR plate, at the same well positions export_plate_layout assigns. Uses real Opentrons labware/pipette API names confirmed against docs.opentrons.com and the Opentrons shared-data labware-definitions repository (opentrons_96_wellplate_200ul_pcr_full_skirt, opentrons_96_tiprack_20ul, opentrons_24_tuberack_nest_1.5ml_snapcap, nest_12_reservoir_15ml, p20_single_gen2) and the confirmed load_labware/load_instrument/transfer method signatures. Master-mix/primer/template/water volumes are clearly-labeled placeholder constants at the top of the script — this is a starting point to review and adapt for your own enzyme and instrument, not a certified ready-to-run protocol.Read-only
export_plate_layoutAssign a set of PCR reactions (name + forward/reverse primer + optional template label) to wells on a 96-well plate, row-major (A1, A2, … A12, then B1, B2, … up to H12). Returns the well-assignment data for rendering a plate diagram; export_opentrons_protocol and export_echo_picklist build their downloadable files from this exact same layout, so all three always agree.Read-only
expression_heatmap_clusterHierarchically cluster a genes x samples expression matrix (UPGMA/average, complete, or single linkage; Euclidean or correlation distance) and return the row/column leaf order, dendrogram merge trees, and row-z-scored values for the Clustered Expression Heatmap visualization.Read-only
fastq_qc_reportFastQC-style deep quality-control report for a FASTQ file: per-base quality and content, GC and length distributions, sequence duplication levels, overrepresented sequences, and adapter content — each with a warn/fail verdict against FastQC's own published thresholds.Read-only
fastq_trimTrim FASTQ reads: an ungapped sliding-suffix adapter match (against the same named Illumina adapters as the QC report) followed by a BWA-style 3' quality trim (the same algorithm Cutadapt's own -q option reuses), then drops reads below a minimum length. Returns the trimmed FASTQ plus before/after read-count, mean-length and mean-quality stats.Read-only
find_orfsFind open reading frames (ATG…stop) across all six frames.Read-only
format_sequenceClean, case-fold, DNA↔RNA convert, reverse and line-wrap a sequence.Read-only
functional_enrichmentOver-representation analysis: test which GO terms (biological process / molecular function / cellular component) and Reactome pathways are statistically enriched in a query gene list versus a background, using the hypergeometric test with Benjamini-Hochberg FDR correction across all tested terms. Uses bundled GO Consortium + Reactome reference data (human only). KEGG is not included (its license does not permit bundling gene sets).Read-only
gc_contentGC content, AT content and per-base composition of a sequence.Read-only
gel_band_sizeCalibrate a gel lane against a marked DNA ladder and read the size of any other band from how far it ran. Migration distance is approximately linear in log(size) over a gel's resolving range, so a least-squares line through the marked rungs interpolates between them — and NOT outside them, which is why a band beyond the marked span comes back flagged `extrapolated` rather than quietly clamped to the nearest rung. Reports R² and the per-rung residual, because one mis-marked rung (a doublet marked as one band, a mis-clicked centre) shifts every size on the gel and a single R² hides it. Distances may be in pixels, millimetres or anything else, as long as they are all measured the same way from the same reference line.Read-only
gene_dossierA gene/drug-target dossier fanned out to five independent sources in one call: Open Targets (function, tractability, top associated diseases), an NCBI/UniProt plain-English function summary, ChEMBL (known drugs and their mechanism/clinical phase, cross-referenced with indications), ClinicalTrials.gov (trials by gene/condition term), and Europe PMC (top cited papers). Each source fails independently — a down source returns null/empty for its own section rather than failing the whole call, and every failure is listed in "sourceErrors" rather than silently omitted.Read-only
gene_expressionA gene's tissue-expression fingerprint: per-tissue median TPM from GTEx (v8) and subcellular localization / RNA tissue-specificity / protein class from the Human Protein Atlas, in one call.Read-only
gene_modelThe real exon/UTR/CDS structure of a human gene's canonical transcript, fetched live from Ensembl (the same exon/CDS map the HGVS Converter tool uses) — for rendering an exon diagram.Read-only
golden_gate_designCHOOSE a set of 4-base Golden Gate/MoClo junction overhangs, rather than scoring one you already have. Maximizes the fidelity of the set's WEAKEST junction against the same published T4-ligase ligation data golden_gate_fidelity scores with, subject to every member actually ligating well — an overhang can score a perfect ratio simply because nothing was ever measured cross-reacting with it, and 94% of the ligation matrix is zeros. Pin the overhangs your vector already commits you to with `fixed`, forbid others with `forbidden`, and give `junctions` when the junctions sit at real positions in real sequence and may only slide a few bases. Reports whether the search was exhaustive (provably the best available) or budget-limited (the best found).Read-only
golden_gate_fidelityScore a candidate set of 4-base Golden Gate/MoClo junction overhangs against real published T4-ligase ligation-count data: per-overhang specificity, the weakest link in the set, and any risky cross-reacting pairs. Optionally compare against a named published overhang set. This is SeqBench's own transparent scoring methodology — it does not reproduce NEB's/Potapov's own published aggregate fidelity percentages for named sets (their exact formula isn't disclosed anywhere accessible).Read-only
golden_gate_from_partsGolden Gate as the reaction runs: digest pre-domesticated part plasmids with a Type IIS enzyme and assemble them in the order their OVERHANGS dictate. The fragment released from each part is the one carrying no recognition site (the site goes out with the backbone, which is why a mis-ordered assembly is not re-cut), and the assembly order is an OUTPUT — a set whose overhangs do not close into a single cycle has no product, and the reason is the answer. Distinct from cloning_simulate's `goldengate` method, which does the other job: designing the primers that ADD the sites to BARE parts, assembled in the order you list them.Read-only
hgvs_convertParse an HGVS "c." variant description (by gene symbol, RefSeq NM_, or Ensembl ENST accession), convert it to genomic (g.) coordinates via a real, live-fetched Ensembl exon/CDS map (transcripts resolved through the bundled MANE RefSeq<->Ensembl crosswalk), apply 3'-rule normalization to any del/dup/ins, and predict the protein (p.) effect where that is safely computable. Also accepts a genomic "g." position on a chromosome or RefSeqGene (NC_/NG_, GRCh38 or GRCh37): it is validated and placed on GRCh38 by NCBI, projected onto the MANE Select transcript by Ensembl, and converted back by this tool's own engine, which must agree — so the answer rests on two independent conversions. And the "NC_/NG_(NM_…):c." form converts on the transcript in parentheses. Refuses cleanly — rather than guessing — for circular/mitochondrial genomes, RNA-level or protein-level input, uncertain/mosaic syntax, splice-junction-adjacent or inversion protein effects, and non-MANE/non-Ensembl transcripts.Read-only
id_map_pollCheck a UniProt id-mapping job submitted via id_map_submit. Returns {status, ready:false} while still running; once FINISHED, also returns the mapped ids (normalized regardless of which target database was requested) and any ids that failed to map.Read-only
id_map_submitSubmit up to 100,000 ids to UniProt's ID mapping service for a single confirmed-safe hop (e.g. Gene_Name -> UniProtKB-Swiss-Prot, or UniProtKB_AC-ID -> Ensembl/GeneID/RefSeq_Protein/Gene_Name). Returns a jobId immediately — poll it with id_map_poll.Changes data
identity_matrixPercent identity between every pair in a multiple alignment. Takes ALIGNED sequences — the output of multiple_sequence_alignment — because inferring an alignment here would bury the aligner's parameters inside a number that reads as a property of the sequences. Two percentages come back per pair: over columns where both have a residue (the usual figure), and over all columns including gaps. The gap between them is the warning — two sequences overlapping in 50 of 2,000 columns read 100% by the first measure and 2.5% by the second.Read-only
in_silico_pcrPredict PCR products for a template and a pair of primers (IUPAC-aware, allows mismatches, handles circular templates). Primers may carry a non-templated 5' tail — a restriction site, a Gibson arm, a Kozak, a tag: a primer primes on its 3' end, and the tail is carried into the product rather than required to match. start/end are the TEMPLATE-derived span, `length` is the whole product including tails, and `features` marks which product bases came from the oligos (present only when there is a tail). Each end reports annealedLength and tailLength.Read-only
kasp_primer_designDesign KASP/ARMS allele-specific genotyping primers for a SNP: two allele-specific forward primers differing only at the 3' terminal base (one per allele), each with the standard KASP universal tail (FAM for allele A, HEX for allele B), a deliberate internal ARMS secondary mismatch near the 3' end whose strength complements that primer's own natural allele mismatch (strong↔weak), and one common downstream reverse primer sized to a chosen amplicon range. Because a forward primer reads the antisense strand, each primer's 3' base sits opposite the complement of the other allele, so the two primers get different mismatch classes and are reported separately (graded from the measured PCR yields in Kwok et al. 1990). Reuses the site's nearest-neighbor Tm engine.Read-only
ligation_setupWork out how many microlitres of vector and insert to pipette to hit a target molar ratio, from each part's length and stock concentration. Handles one insert or several with independent equivalents (Gibson, Golden Gate, MoClo), reports pmol and ng per part alongside the volumes, and flags the two things that actually go wrong on a bench: a volume below what a pipette measures reliably, and a plan whose DNA does not leave room for buffer and enzyme. A molar ratio is about moles, so a shorter insert at 3 molar equivalents goes in at LESS mass than the vector — that conversion is the point.Read-only
melting_temperaturePrimer/oligo melting temperature: nearest-neighbor (SantaLucia 1998) at the supplied reaction conditions, recommended from 14 nt up, with the Wallace rule for shorter oligos, a fixed-100 mM-Na+ Schildkraut-Lifson reference estimate, and molecular weights.Read-only
motif_finderFind (overlapping) occurrences of an IUPAC motif on either strand, allowing mismatches.Read-only
multiple_sequence_alignmentCenter-star multiple sequence alignment of a multi-FASTA input — nucleotide or protein, detected from the records and reported as `type` — with consensus and per-column conservation.Read-only
multiplex_panel_designChoose one primer pair per target so the whole panel works in one tube: no cross-dimer between any two of the primers, every amplicon resolvable from every other on the gel you will run, and one annealing temperature that serves all of them. Searches combinations rather than picking each target's best pair in isolation, which is what makes panels fail — and when no compatible panel exists it names the target pairs that cannot be multiplexed at all, so you know which one to redesign.Read-only
nonrepetitive_parts_designBuild a set of new genetic parts that match a degenerate (IUPAC) template and share more than a chosen length with nothing — not each other, not themselves, not a background sequence you supply. This is how a toolbox of promoters, RBSs, terminators or sgRNA handles is made large without making an assembly unstable. Honours a GC range and excluded motifs (restriction sites) during the search rather than filtering afterwards. Deterministic given a seed: the same inputs give the same toolbox. Returns fewer parts than asked, with a reason, when the constraints leave no room — it never invents a repetitive one to hit the count.Read-only
nonrepetitive_parts_findGiven a toolbox of genetic parts, return the largest subset in which no two parts share more than a chosen length of sequence — on either strand. Parts that share a long stretch recombine into each other in a multi-part assembly and are the single largest cause of DNA synthesis failure, and neither shows up when the parts are checked one at a time. Reports every conflicting pair and why each dropped part was dropped. Deterministic: no model, no rate, no score. Use nonrepetitive_parts_design to build new parts instead of selecting from existing ones.Read-only
od600_cellsConvert an OD600 reading to a cell density and a total cell count, correcting for the optical path and any dilution, and flagging the two things that silently make the number wrong: a reading above the linear range (a dense culture reads LOW), and a plate-reader path length assumed to be a cuvette's 1 cm. Also plans a back-dilution to a target OD. The cells-per-OD factor is a convention, not a constant — it varies with strain, growth phase and instrument, so pass factorPerOd when you have calibrated against plate counts.Read-only
oligo_analysisFull oligo analysis: nearest-neighbor Tm/ΔG/ΔH/ΔS plus hairpin and self-dimer screening with base-pair diagrams and warnings.Read-only
oligo_cofoldMinimum-free-energy structure and ΔG for one oligo (hairpin) or two oligos together (homo/heterodimer), using ViennaRNA's published loop model at a temperature you choose — DNA parameters (Mathews 2004) by default, RNA (Turner 2004) on request. Reports each strand alone, the duplex, and the interaction ΔG the two gain by pairing with each other rather than folding alone, which is the number a primer-dimer screen wants. Unlike oligo_analysis's fast stack-sum screen this is a full loop model with bulge, internal-loop and dangling-end terms; the two are on different parameter sets and must not be compared. PREDICTED, NOT MEASURED. No skill statistic is claimed for predicting whether a PCR fails. Loop-model MFE folding reproduces measured structure well for short duplexes and progressively worse with length; the ΔG itself carries roughly kcal/mol-scale uncertainty and the MFE structure is one structure out of an ensemble — request `partition` for the ensemble free energy, which is the more honest single number when several structures compete. Valid for: short oligos, at most 200 nt per strand, at the temperature given. It models two strands in isolation at no particular concentration: it does not know your primer concentration, salt, or cycling program, so it cannot say whether a dimer will actually form in your tube.Read-only
oligo_pool_screenScreen a whole set of oligos you already have — every pair for cross-dimers, every oligo for its own hairpin and self-dimer, and the set for duplicates and Tm spread — and get back the conflicts ranked rather than a table of every combination. This is the pool-level answer cross_dimer gives one pair at a time: 51 primers is 1,275 pairs, which is 1,275 separate calls done by hand and one call done here. Not to be confused with multiplex_panel_design, which DESIGNS primers from templates; this takes the primers you have already ordered. A pairing that involves an oligo's 3' END is judged at a weaker ΔG than one that only pairs internally, because that end is where extension starts — the same two-bar rule the multiplex panel designer uses. Every number is a nearest-neighbor calculation over the sequences supplied, not a prediction of what the reaction will do.Read-only
operon_designBuild a polycistronic operon from a promoter, a list of CDSs with their RBSs, spacers and a terminator — optionally recoding every CDS for a host — then scan the ASSEMBLED molecule for internal promoters, Shine-Dalgarno sequences, terminators, out-of-frame start codons and repeats. Scanning the product rather than the parts is the point: these elements are very often created BY THE JOIN between two parts, so checking each part alone finds nothing. Returns the sequence, an annotated GenBank file, and every element found. Assigns no translation rate — use rbs_library_design to choose RBSs, then rbs_predict to rank the result.Read-only
operon_scanScan a multi-gene construct, on both strands, for the sequence that quietly breaks operons: promoter-like -35/-10 pairs (including ones pointing backwards, which make antisense RNA), Shine-Dalgarno sequences positioned in front of an internal start codon, terminator-shaped hairpins with a U-tract, out-of-frame start codons inside a declared CDS, and exact direct repeats. Reports what MATCHED and how far it sits from consensus — it does not score a match or claim it transcribes. For an estimated promoter strength use promoter_predict; for a translation rate use rbs_predict.Read-only
ortholog_mapLook up the orthologous (or paralogous) gene for up to 50 gene symbols in a target species, via Ensembl's homology-by-symbol REST endpoint. Symbols with no homology record are reported in `unmapped`, never silently dropped.Read-only
outcome_deconvolveDecompose one Sanger trace into fractions over a set of candidate molecules — the intended construct, the empty backbone, a double insert, a flipped part — instead of onto a generic indel ladder. Non-negative least squares against the candidates' own sequences, so no molecule is ever assigned a negative share. Reports the R² of the fit, which is what says whether the tube holds anything outside the candidate set, and GROUPS candidates the read cannot tell apart rather than splitting their share between them. Feed it the alternatives from assembly_outcomes. PREDICTED, NOT MEASURED. Every run reports its own R²: how much of the observed per-position composition the candidate basis explains. That is a measured adequacy on YOUR trace, and a low value is a statement that the tube holds something the candidate set does not contain. No accuracy against a reference method is published for this implementation, and none is quoted. Exact recovery of synthetic mixtures is deliberately NOT offered as validation: it tests the arithmetic, not whether a real capillary trace behaves like the model. Valid for: One read, from one primer, over a set of candidate molecules that all contain that primer's site and differ from each other within the read. Fractions are reported per GROUP of candidates the read cannot tell apart, and that grouping is part of the answer. NOT valid when the read's anchor identity to a candidate is low (it does not share the primer region), nor when R² comes back low, nor for telling apart candidates that differ only outside the read.Read-only
pairwise_alignmentGlobal (Needleman-Wunsch), local (Smith-Waterman) or semi-global/fitting pairwise alignment of two sequences, with match/mismatch scoring and affine gap costs (Gotoh).Read-only
parse_genbankParse a GenBank flat file into its locus, definition, features and sequence. An EMBL / ENA flat file is accepted too — the same INSDC record in a different layout — and reported with sourceFormat "embl".Read-only
parse_gff3Read a GFF3 annotation file: every feature with its 1-based inclusive coordinates, strand, phase, score and decoded attributes, plus any sequences from an embedded ##FASTA section. GFF3 usually carries NO sequence — it annotates a separate FASTA, joined on the first column — so the result says which of the two it was handed. Give `sequence` (or let an embedded ##FASTA supply it) to also get the record as GenBank, readable by every other tool here. Percent-escapes are decoded, `.` is read as absent rather than zero, and a feature repeated across lines keeps both lines.Read-only
parse_sanger_traceDecode a Sanger ABIF (.ab1 / .abi) chromatogram: base calls, per-base quality, the four dye-channel traces, peak locations, and the run's own labels (sample name, well, plate, instrument, run start).Read-only
parse_snapgeneRead a SnapGene .dna file: sequence, topology, every feature with its span, strand, display color and qualifiers (spliced and origin-spanning features kept as such), and the saved primer list.Read-only
parts_library_searchSearch a parts list harvested from the annotated features of the vector library — promoters, terminators, RBSs, polyA signals, origins, selection markers, affinity tags, reporters, linkers/MCSs — by name, kind or length. Nothing here is transcribed: every part is the exact sequence a GenBank record annotated, and each hit carries the accession and 1-based span it was cut from, plus every other library vector the same part was found in. Parts whose location is spliced or approximate are excluded, because their sequence is not fully determined.Read-only
plasmid_annotateAuto-detect common cloning features (promoters, tags, origins, resistance markers, MCS, primers) on both strands. Signatures under 20 bp must match exactly; longer ones tolerate up to ~10% mismatches so point mutants still annotate — each feature reports its own `mismatches` count and an `exact` flag.Read-only
plasmid_deep_annotateAnnotate a plasmid against pLannotate's open-source feature library — a much larger signature set (GenoLIB parts + Swiss-Prot + FPbase + Rfam, cross-referenced against ~195k Addgene-deposited plasmids) than plasmid_annotate's built-in curated list, and it reports partial and low-identity hits as graded alignments rather than the pass/fail signature match plasmid_annotate does (that one is not exact-only either — signatures of 20 bp or more tolerate up to ~10% mismatches — but it reports a hit or nothing, with a `mismatches` count and an `exact` flag). Each feature here carries its percent identity, reference coverage and a fragment flag so you can judge a weak hit. Runs a multi-second search on a shared service and is therefore rate limited (see 429/503); use plasmid_annotate for an instant, unmetered first pass.Read-only
plasmid_dossierAnswer the question someone holding an unlabelled tube actually has, rather than returning another feature list: which antibiotic to select on, which host will carry it and at roughly what copy number, which curated backbone it resembles, and what on it will bite you later (duplicate markers, two origins for the same host, internal Type IIS sites that stop it being a Golden Gate part, long direct repeats a recA+ host can recombine out). Reads features from an annotated GenBank record when you paste one, otherwise runs the built-in signature scan. EVERY ABSENCE TRAVELS WITH THE VOCABULARY IT WAS DECIDED AGAINST: an empty marker list means nothing matched the families this screen knows, never that the plasmid cannot be selected — run plasmid_deep_annotate, whose library is far larger, before concluding that.Read-only
plasmid_full_reportOne combined view of 'what is this plasmid': recognized common features (from plasmid_annotate), backbone identity / possible chimera (from plasmid_identify), and — the two crossed together — any region that neither a curated backbone nor a recognized common feature explains. That last list is a triage signal (an unusual insert, an unannotated part, or worth a closer look), not a defect finding: a real gene-of-interest legitimately has no curated-feature match.Read-only
plasmid_identifyScreen a query plasmid against a small curated set of common backbones (cloning vectors, expression vectors, BACs — see referencesChecked for the exact list) to identify which one(s) it resembles, separate an unmatched region (normal — your own insert) from a POSSIBLE CHIMERA (a region matching a different known backbone than its neighbor), and report per-match %identity/%coverage. NOT a search against Addgene's ~100k-plasmid catalog or PlasmidScope's 850k+ — a curated-set screen only.Read-only
prime_editing_designDesign SpCas9 prime-editing pegRNAs for a substitution, insertion, deletion, or small replacement: for each usable NGG PAM it builds the spacer, a primer-binding-site (PBS) length sweep targeting a ~30 C melting temperature, the reverse-transcriptase template (RTT) that encodes the edit, and the full 3' extension, plus PE3 nicking-sgRNA suggestions 40-90 bp away on the opposite strand. Designs where the edit destroys the pegRNA's own PAM (preventing re-nicking of the edited allele) are ranked first. Coordinates: every pegRNA coordinate (protospacer span, nick position, editStart/editEnd) is 1-based inclusive in the submitted PRE-EDIT target's frame — the protospacer+PAM search runs on the unedited sequence, because Cas9 has to bind the allele you actually have. The one exception is edit-dependent PE3b nicking guides, which exist only once the edit is installed; each nickingGuides entry therefore carries a `coordinateFrame` field of "target" or "editedSequence" naming the frame its own start/end/nickToNickDistance are measured in, and for a length-changing edit the two frames differ downstream of the edit. Off-target activity is not evaluated (no in-browser reference genome).Read-only
prime_editing_efficiencyPredict per-pegRNA prime-editing efficiency for one edit with PRIDICT2.0, and return the top-scoring pegRNA designs ranked by it. Takes the target as context, the edit in brackets, then context — ACGT...(A/G)...ACGT, with roughly 100+ bp each side — and enumerates PBS/RTT length combinations, scoring every one in HEK293 and K562. Each candidate comes back with both scores, its percentile against the training library, its rank, the spacer, PBS and RTT lengths, the full pegRNA, and Golden Gate cloning oligos. Use it to CHOOSE between designs; the number is not a promised editing percentage. PREDICTED, NOT MEASURED (Spearman ρ = 0.85 on held-out data from the libraries it was trained on). Spearman rho of about 0.85 for intended edits on held-out library data — the best-validated figure of any model in this registry, and roughly double OSTIR's 0.39 on independent data. That figure is still within the library and cell lines it was trained on. Valid for: human sequence, and efficiency ranking within one locus. It is parameterized on HEK293 and K562; your cell type, delivery method, and chromatin context will all move the absolute efficiency, chromatin alone by severalfold. Nothing here is predicted for a non-human host or for editors outside the PE2/PE3 architecture the training libraries used.Read-only
prime_editing_twin_designDesign a twinPE pegRNA pair (Anzalone et al. 2022) for a replacement too large for a single pegRNA's RTT: a left pegRNA nicks the + strand at/before the replacement window and a right pegRNA nicks the - strand at/after it, each synthesizing a new 3' flap; both flaps are truncated at a shared overlap in the middle of the new sequence so they anneal and resolve the edit without an HDR donor. Coordinates: both pegRNAs' protospacerStart/protospacerEnd/nickPosition are 1-based inclusive in the submitted PRE-EDIT target's frame (the PAM search runs on the unedited sequence, on both sides), while replaceSpan is the span of the new content in the returned editedSequence. Off-target activity is not evaluated (no in-browser reference genome).Read-only
primer_designDe-novo PCR primer design (Primer3-style penalty picker): enumerate and score candidate primer pairs against length/Tm/GC/3'-clamp/structure constraints.Read-only
primer_site_accessibilityFold the template around each place a primer binds, at the annealing temperature and under DNA parameters, and report how much of the primer's own footprint sits inside a helix. Primer design tools score the oligo — its Tm, its hairpin, its dimers — and leave the template unexamined, while a binding site buried in a stable stem is an ordinary reason a well-designed primer does not amplify. Returns the folded window, its free energy, the paired fraction of the footprint, and how many of the 3'-terminal five bases are paired.Read-only
primer_specificitySelf-hosted e-PCR-style screen for off-target amplicons predicted by a primer pair against a small set of curated reference genomes (currently: E. coli K-12 MG1655, B. subtilis 168, human mitochondrion rCRS, Mycoplasma hyorhinis SK76 — see genomesChecked in the response for the exact list, and note that the nuclear human and mouse genomes are NOT covered). Amplicons are 1-based inclusive on the plus strand; a product across a circular genome's origin reports an end lower than its start and sets wraps: true. This checks background/host-genome specificity, NOT whether the primers hit your intended target — pair it with in_silico_pcr against your own template for that. Each off-target end reports its 3' ANCHOR — the primer's unbroken run of matched bases at the extending end — with that anchor's nearest-neighbor ΔG and a margin against the intended, fully matched reaction, so a site can be told apart by WHERE its mismatches fall rather than only how many there are: one mismatch at the 5' end leaves a site nearly as strong, and one at the 3' base leaves it unable to prime at all. Batchable over candidate REVERSE primers against one fixed forward primer (screen many candidates against a shared partner) — not independent primer-pair batching, which this tool doesn't support. A primer may carry a non-templated 5' tail (a restriction site, a Gibson arm, a tag): the screen looks for a 3'-anchored annealing region as well as a full-length match, so a tailed cloning primer is screened rather than silently matching nothing. Each end's `anchor` is the annealed run, which is the length that matters for extension, and `start`/`end` are measured on the ANNEALED footprints — the bases each primer actually pairs with on the genome — so `length` (the product, tails included) equals end - start + 1 only for untailed primers. Screening a TAILED primer without `intendedTemplate` inflates every margin by the tail's own free energy, because nothing about an oligo says where its non-templated part ends; `intended.basis` reports which footprint the margins rest on.Read-only
promoter_library_designBuild a set of sigma-70 promoters that (a) share no more than a chosen length of sequence with each other, so the library does not recombine with itself once integrated, and (b) span a range of predicted transcription rates, picked as an evenly log-spaced ladder. Variants are constructed deterministically from an IUPAC template holding the -35 and -10 consensus; their strengths are then estimated by the Promoter Calculator model. Rungs with no variant near them are reported as gaps rather than filled with the nearest thing. PREDICTED, NOT MEASURED: strengths carry R^2 = 0.45-0.60 against independent in vivo data, so treat the ladder as a ranked set to screen, not as calibrated numbers. PREDICTED, NOT MEASURED. R^2 = 0.45 and 0.60 against the two INDEPENDENT in vivo datasets the authors tested (Hossain et al., 4,350 promoters, Spearman rho = 0.69; Urtecho et al., 10,898 promoters, rho = 0.67). The widely quoted R^2 = 0.80 is a held-out tenth of the authors' OWN in vitro transcription data and is not the number to plan against: a promoter in a cell is the in vivo case, where between a third and a half of the variance is unexplained. Valid for: sigma-70 (housekeeping) promoters in E. coli. NOT valid for another sigma factor, another organism, a promoter under activator or repressor control, or anything about mRNA stability or translation — rbs_predict is the translation half, and neither speaks to the other.Read-only
promoter_predictScan DNA for E. coli sigma-70 promoters and estimate each one's transcription initiation rate, with the free-energy terms it is built from: the -35 and -10 boxes, the spacer, the discriminator, the extended -10 and the initial transcribed region. Pairs with rbs_predict — together they separate 'nothing is transcribed' from 'it is transcribed and not translated', which no single measurement on the sequence does. Both strands by default, because a promoter reading into your insert from the other strand is still a promoter. PREDICTED, NOT MEASURED. R^2 = 0.45 and 0.60 against the two INDEPENDENT in vivo datasets the authors tested (Hossain et al., 4,350 promoters, Spearman rho = 0.69; Urtecho et al., 10,898 promoters, rho = 0.67). The widely quoted R^2 = 0.80 is a held-out tenth of the authors' OWN in vitro transcription data and is not the number to plan against: a promoter in a cell is the in vivo case, where between a third and a half of the variance is unexplained. Valid for: sigma-70 (housekeeping) promoters in E. coli. NOT valid for another sigma factor, another organism, a promoter under activator or repressor control, or anything about mRNA stability or translation — rbs_predict is the translation half, and neither speaks to the other.Read-only
protease_digestionIn-silico protease/chemical digestion: cleave a protein and report each peptide's position, length and neutral mass. CNBr masses assume terminal Met becomes homoserine lactone; the sequence retains M and the response labels the modification.Read-only
protein_annotate_pollCheck an InterProScan job submitted via protein_annotate_submit. Returns {status, ready:false} while still running; once FINISHED, also returns the parsed domain architecture, per-match details and deduplicated GO terms.Read-only
protein_annotate_submitSubmit a protein sequence to EBI InterProScan for domain architecture, family and GO-term annotation. Returns a jobId immediately — the job itself takes minutes; poll it with protein_annotate_poll.Changes data
protein_hydrophobicitySliding-window hydropathy/hydrophobicity profile (ProtScale-style) over a published amino-acid scale.Read-only
protein_propertiesProtein properties: molecular weight, isoelectric point, GRAVY, extinction coefficient and composition.Read-only
qpcr_ddctRelative quantification from Ct values: ΔCt, ΔΔCt and fold change, with the standard error propagated from replicates. Picks the method from the amplicon efficiencies rather than leaving it to you — Livak's 2^-ΔΔCt when both are 100% efficient, the Pfaffl efficiency-corrected ratio when they are not, since running Livak over amplicons that do not qualify is the failure this tool exists to prevent. Reports what Livak would have said, so the size of that difference is visible.Read-only
random_sequenceGenerate a random DNA, RNA or protein sequence, optionally with a target GC content.Read-only
rbs_designDesign a 5' UTR / ribosome binding site for a given CDS. Generates a spread of Shine-Dalgarno cores and SD-to-start spacings, scores every one with OSTIR in the context of your own CDS (which matters — the rate depends on how the RBS interacts with that CDS's 5' folding), and returns them ranked. Supply targetExpression to rank by closeness to a target rate instead of by maximum strength, and supply your existing 5' UTR to get a measured baseline and fold-change for each candidate. Runs ViennaRNA on a shared service and is therefore rate limited (see 429/503). PREDICTED, NOT MEASURED (Spearman ρ = 0.39 on two 5' UTR datasets it was not fitted to). Spearman ρ = 0.39 against measured expression on two 5' UTR datasets it was not fitted to (Gilliot & Gorochowski, Nucleic Acids Res 2024;52(13):e58); reproduced in this repo at ρ = 0.509 on an independent leave-one-context-out split over 194,636 measurements, and at a mean ρ = 0.62 over fifteen further measurement systems it was not fitted to — six species, four reporters, three readouts — on which every learned model tried here scored BELOW it (scripts/rbs-eval) — above the published figure because 65,841 rows of that compilation record no measurement at all (fluorescence mean and s.d. both exactly 0) and are excluded here rather than scored as the weakest expression observed. The widely quoted 53% within 2-fold / 91% within 10-fold are calibration residuals on the fitting set, not held-out validation. Valid for: translation INITIATION only, in E. coli-like Gram-negative hosts (the model is parameterized on the E. coli anti-Shine-Dalgarno sequence). Rankings within one construct context; the absolute value has no units and no meaning.Read-only
rbs_library_designBuild a ladder of ribosome binding sites whose predicted translation initiation rates are evenly spread, in log space, across a range you choose — the standard way to titrate one enzyme's level in a pathway without guessing. Scores every SD-core/spacing/spacer-composition combination with OSTIR in your own CDS context, then picks one variant per rung. Rungs it cannot fill are reported as GAPS rather than filled with the nearest available variant, so a library that does not really span the range says so. PREDICTED, NOT MEASURED: the ordering comes from a model with ρ ≈ 0.39 against measured expression, which is what makes it usable for ranking a library you will screen and unusable for hitting an absolute number. PREDICTED, NOT MEASURED (Spearman ρ = 0.39 on two 5' UTR datasets it was not fitted to). Spearman ρ = 0.39 against measured expression on two 5' UTR datasets it was not fitted to (Gilliot & Gorochowski, Nucleic Acids Res 2024;52(13):e58); reproduced in this repo at ρ = 0.509 on an independent leave-one-context-out split over 194,636 measurements, and at a mean ρ = 0.62 over fifteen further measurement systems it was not fitted to — six species, four reporters, three readouts — on which every learned model tried here scored BELOW it (scripts/rbs-eval) — above the published figure because 65,841 rows of that compilation record no measurement at all (fluorescence mean and s.d. both exactly 0) and are excluded here rather than scored as the weakest expression observed. The widely quoted 53% within 2-fold / 91% within 10-fold are calibration residuals on the fitting set, not held-out validation. Valid for: translation INITIATION only, in E. coli-like Gram-negative hosts (the model is parameterized on the E. coli anti-Shine-Dalgarno sequence). Rankings within one construct context; the absolute value has no units and no meaning.Read-only
rbs_occlusionFind every stretch of a transcript that is reverse-complementary to the ribosome's footprint (-20 to +13 around the start codon), with the length and nearest-neighbour Tm of each duplex, and flag the ones that cover the Shine-Dalgarno core or the start codon. This is the mechanism every translational switch runs on — riboswitch, toehold switch, RNA thermometer, antisense repressor — and the first thing to look at when a switch does not switch or a construct is unexpectedly silent. Also reports where a sensor domain can be inserted without touching the site. Deterministic base-pairing, not a folding prediction: use rna_fold to ask whether a given stem actually wins at 37 °C. Reports no switching ratio — see this tool's notes for why.Read-only
rbs_predictPredict the translation initiation rate at each start codon in a bacterial mRNA using OSTIR, the open-source continuation of the Salis lab RBS Calculator, with ViennaRNA free energies. Returns the predicted rate plus the full thermodynamic breakdown (16S rRNA:mRNA hybridization, mRNA unfolding, spacing, standby site, start-codon binding) for every start codon found. Rates are on an arbitrary scale — compare them as ratios, not as absolute expression levels. Runs ViennaRNA on a shared service and is therefore rate limited (see 429/503). PREDICTED, NOT MEASURED (Spearman ρ = 0.39 on two 5' UTR datasets it was not fitted to). Spearman ρ = 0.39 against measured expression on two 5' UTR datasets it was not fitted to (Gilliot & Gorochowski, Nucleic Acids Res 2024;52(13):e58); reproduced in this repo at ρ = 0.509 on an independent leave-one-context-out split over 194,636 measurements, and at a mean ρ = 0.62 over fifteen further measurement systems it was not fitted to — six species, four reporters, three readouts — on which every learned model tried here scored BELOW it (scripts/rbs-eval) — above the published figure because 65,841 rows of that compilation record no measurement at all (fluorescence mean and s.d. both exactly 0) and are excluded here rather than scored as the weakest expression observed. The widely quoted 53% within 2-fold / 91% within 10-fold are calibration residuals on the fitting set, not held-out validation. Valid for: translation INITIATION only, in E. coli-like Gram-negative hosts (the model is parameterized on the E. coli anti-Shine-Dalgarno sequence). Rankings within one construct context; the absolute value has no units and no meaning.Read-only
read_placement_planGiven the molecules a cloning reaction could have produced and the sequencing primers you could use, work out which primers separate which pairs of candidates — and return the smallest set that separates every pair any of them can. Scores a read by discrimination, not coverage: a read of 700 bases every candidate shares is worth nothing, and a short read across a junction is worth everything. Names the pairs no primer here separates, so you find out before paying for the reads rather than after. Pair it with outcome_deconvolve once the traces come back.Read-only
repeat_instabilityFind the exact direct repeats in a construct that make it deletable, and build the molecule each pair would collapse to. Two copies of the same terminator or promoter in a multi-gene assembly let the DNA between them recombine out — silently, so the clone grows and the map looks right until it is sequenced. Returns each repeat pair's coordinates plus the resulting sequence(s), ordered by repeat length and spacer, the two factors that govern how readily a pair recombines. Reports no deletion RATE: none is derivable from sequence alone. Feed a deletion product to diagnostic_digest to screen for it.Read-only
restriction_sitesFind restriction enzyme recognition sites in a DNA sequence.Read-only
reverse_complementReverse, complement and reverse complement of a DNA or RNA sequence.Read-only
reverse_translateBack-translate a protein to DNA (most-frequent codon per organism, or degenerate IUPAC consensus).Read-only
riboswitch_statesFold a transcript twice — once freely, once with the aptamer held out of the secondary structure as a bound ligand would hold it — and report what changes at the ribosome binding site. Returns both structures, both free energies, the cost of occupying the aptamer, and whether the site becomes more or less accessible. This is the question a riboswitch design can be checked on before the bench: is the site sequestered in one state and free in the other. Returns NO activation ratio and no fold-change — see this tool's validation note for why that number is not offered. Pair with rbs_occlusion, which finds what can pair over the site in the first place. PREDICTED, NOT MEASURED. None is published for this combination, and none is claimed. The two folds are ViennaRNA MFE structures with its documented parameter set; what is NOT established is that the ligand-bound state is modelled correctly by holding the aptamer out of the secondary structure, which is an approximation of a three-dimensional binding event. No activation ratio is returned, because the only published validation set for predicting one (Espah Borujeni et al., NAR 2016;44(1):1) is CC BY-NC and unusable here. Valid for: single-strand RNA up to 1,000 nt at one temperature, with no pseudoknots and no tertiary structure. A translational riboswitch whose mechanism is sequestering the ribosome binding site — NOT transcriptional attenuators, NOT ribozymes, and NOT any switch whose aptamer overlaps the site it regulates.Read-only
rna_foldPredict an RNA secondary structure by minimum free energy (MFE) with ViennaRNA's full Turner 2004 loop model at 37 °C (no pseudoknots). Returns the dot-bracket structure, the MFE (kcal/mol), the list of base pairs, and the engine that produced them. PREDICTED, NOT MEASURED. Benchmarks of MFE folding against known structures put single-sequence accuracy near 40-70% of base pairs recovered, falling as the sequence gets longer; no held-out statistic is measured here. The MFE structure is one structure out of an ensemble, and a second structure within a fraction of a kcal/mol of it is common. Valid for: a single strand up to 2,000 nt at a fixed 37 °C (600 nt when folded in the workbench, which runs the fallback engine in the browser). No pseudoknots, no two-strand hybridization — use oligo_cofold for that — and no temperature dependence, so a structure predicted here is not a structure at your annealing temperature.Read-only
sanger_assembleAssemble two or more Sanger reads into a contig WITHOUT a reference sequence — the forward/reverse pair of one insert, or a set of tiling reads. Orientation is worked out from the overlaps, ends are quality-trimmed, and the consensus is quality-weighted. Reports depth and agreement per position, every position where the reads disagree (including a base only one read has), and any read that overlapped nothing. Use sanger_vs_reference instead when you already know what the sequence should be.Read-only
sanger_indel_spectrumQuantify CRISPR editing from a pair of Sanger traces — an unedited control and the edited pool — by decomposing the edited trace onto shifted copies of the control. Returns the indel spectrum (how much of the pool carries each insertion or deletion size), the unedited fraction, and the R² of the decomposition, which is the number that says whether the model fits your traces at all. Non-negative least squares, so no allele is ever assigned a negative share. Does not work for base editing, which makes a mixed base rather than a shift. PREDICTED, NOT MEASURED. Every run reports its own R²: how much of the observed window the shifted-control basis actually explains. That is a measured adequacy of the model on YOUR traces, and a low value means the assumption is wrong here rather than that the edit is weak. On accuracy against a reference method for real samples, none is published for this implementation — the underlying decomposition is TIDE's, whose authors report their own concordance with amplicon sequencing, and that number does not transfer to this code so it is not quoted. This implementation's exact recovery of synthetic mixtures is deliberately not offered as validation either: it tests the arithmetic, not whether the model fits a real trace. Valid for: A pool of alleles that differ from one control read by simple insertions or deletions at a known cut site, where both reads come from the same amplicon and chemistry and both extend well past the cut. NOT valid for substitution-only edits — base editing produces a mixed base, not a shift, and this model cannot see it — nor for a knock-in whose insert is novel sequence rather than a frame shift of the control, nor for any run whose R² comes back low.Read-only
sanger_knockin_quantMeasure the rate of a SPECIFIC intended edit from a pair of Sanger traces — an unedited control and the edited pool — by decomposing the edited trace onto three things at once: the wild-type allele, the intended edited allele, and the unintended indels. Serves both readouts that need this: HDR knock-in rate (what fraction of the pool carries the donor's edit, including an insert of novel sequence), and prime editing (the pegRNA's intended substitution, insertion, deletion or replacement as the intended column, and the indel byproducts at the nick as the shift columns). This is what sanger_indel_spectrum cannot do: that tool's basis is indexed by indel LENGTH, so an intended 6 bp knock-in and an accidental 6 bp NHEJ deletion are one column there. Returns knock-in / wild-type / unintended-indel percentages, the byproduct spectrum by shift, and the R² that says whether the model fits your traces at all. Non-negative least squares, so no allele is ever assigned a negative share. For a substitution or replacement the reference allele you name is checked against the control read before anything is fitted; an insertion and a deletion name no reference bases, so there only the position can be range-checked. PREDICTED, NOT MEASURED. Every run reports its own R²: how much of the observed window the basis actually explains, measured on YOUR traces, where a low value means the model is wrong here rather than that the edit is weak. On agreement with a reference method for real samples — amplicon sequencing or clonal genotyping — none is published for this implementation. The underlying decomposition is TIDE/TIDER's, whose authors report their own concordance, and that number does not transfer to this code so it is not quoted. This implementation's near-exact recovery of synthetic mixtures is deliberately not offered as validation: a synthetic mixture is built from the same idealised one-hot peaks the basis assumes, so recovering it tests the arithmetic and cannot test the assumption. Valid for: A pool whose intended edit is known EXACTLY, read against a control amplicon of the same locus and chemistry, with both reads extending well past the edit. The novel inserted bases of a knock-in carry an assumed peak shape rather than a measured one (returned as constructedPositions) — the more of the window they occupy, the more of the fit is testing that assumption. NOT valid when the reported R² is low; nor for separating an intended pure DELETION from an unintended indel of the same net length ANYWHERE in the window, not only one at the same site (the tool reports which case it is in `sameShiftByproduct`: when that column is not fitted, knockinPercent is the sum of the two); nor for telling an on-target knock-in from a random integration of the same donor; nor for resolving haplotypes, since a Sanger trace of a pool has no phase information.Read-only
sanger_plate_verifyJudge a whole plate of Sanger reads against one construct and return one row per clone: PASS, POINT_MUTATION, INDEL, VECTOR_ONLY (the insert is absent), WRONG_INSERT (the backbone matches and the insert does not), LOW_COVERAGE, or AMBIGUOUS. Reads are grouped into clones from their FASTA/FASTQ record names (facility conventions like PlateA_A01_pXY-1_M13F, pXY-1_T7-F, 2026-08-01_pXY_clone3_R), and every read's assignment is reported with a confidence so a grouping can be corrected rather than trusted. Each clone's reads are piled up in reference coordinates, so a difference one read reports where other covering reads read the reference is reported as the sequencing error it is, not as a mutation — and a position no read covered is never PASS. Every verdict cites the positions it rests on. Give insertStart/insertEnd to have clones judged over the insert alone, which is also what VECTOR_ONLY and WRONG_INSERT need.Read-only
sanger_vs_referenceAlign a Sanger ABIF read to a reference and report identity plus every mismatch, insertion and deletion.Read-only
save_permalinkRun a registered tool and save its (arguments, result) pair under a short permanent code that anyone with the link can view read-only (/permalink/{code}). Use this to cite or share a specific result (e.g. a verify_construct or verify_assembly check) rather than re-pasting it.Destructive
seqfile_statsStatistics for a FASTA or FASTQ file: count, length distribution, N50, GC content and (FASTQ) mean quality.Read-only
sequence_fetchFetch a public DNA/protein record by accession from NCBI Nucleotide, NCBI Protein, UniProt, or Ensembl (e.g. NM_000546, NP_000537, P04637, ENSG00000141510). Only the accession is sent upstream. Use sequence_search first if you only know a gene/organism name, not an accession. For an Ensembl transcript ID this returns spliced cDNA; for a gene ID it returns the full genomic locus (introns included) — Ensembl's own default for each ID type.Read-only
sequence_format_convertConvert between FASTA and GenBank (whole sequence, CDS or protein), or export to TSV.Read-only
sequence_reportOne-click DNA analysis: composition, ORFs, restriction-enzyme scan (single cutters) and end-primer Tm composed into a single report with a copyable text block.Read-only
sequence_searchResolve a gene/organism name — or a raw NCBI search term — to candidate accessions, instead of guessing one. Returns up to maxResults hits (accession, title, organism); pass the accession you want to sequence_fetch.Read-only
sequencing_readback_verifyAlign raw Sanger or NGS reads (FASTA or FASTQ) back onto a claimed reference sequence using minimap2, and report per-read mapping identity plus exact variant positions (substitutions/insertions/deletions), with a consensus view across reads and a corrected consensus sequence (the reference with every consensus-supported edit applied). Each alignment also reports how much of the READ was used (queryCoveragePct/clippedBases), since identity is measured over the aligned portion only and a partially-used read would otherwise score perfectly. Set circular: true for a plasmid so reads crossing the reference's arbitrary linear start are aligned through the join rather than cut short at it. Also calls STRUCTURAL variants from split alignments — a large deletion, tandem duplication, inversion or backbone rearrangement never appears as a run of mismatches, only as one read aligning at several distant reference positions, so per-base calling reports a perfect clone — and returns a coverage depth profile with the regions no read reached at all, since "never read" is not "correct". On a circular reference one junction cannot always tell an event of length d from one of length referenceLength − d the other way round; where the read's own blocks and the coverage profile settle it they do, and where they do not the call carries an alternateInterpretation with the other reading rather than presenting one as a finding. Set platform (nanopore/pacbio/illumina/sanger) to pick minimap2's preset; the preset used is reported back. Complements verify_construct/verify_assembly: those re-derive what a design SHOULD produce from its own stated inputs; this checks what a real sequencer actually read back.Read-only
session_createStart a scratch session that holds several named sequences/values (e.g. vector, insert, forward/reverse primer) for use across multiple tool calls via session_run, instead of re-pasting them into every call. Sessions expire after 24 hours.Changes data
session_getFetch named entries from a session. Prefer session_run for actually USING the values — it keeps raw sequences out of your context. Use this mainly to inspect or debug what a session currently holds.Read-only
session_runRun any SeqBench tool, resolving selected arguments from a session's named entries instead of pasting them inline, and optionally store selected result fields back into the session by name. This is the main way to chain a multi-part design (vector + insert + primers) across calls without shuttling raw sequences through your own context.Destructive
session_setAdd or overwrite named entries in an existing session.Destructive
sirna_designDesign siRNA duplexes against an mRNA target using the established Reynolds (2004) 8-criteria score and the Ui-Tei (2004) rules, plus the siDirect seed-duplex Tm off-target flag (≥21.5 °C, computed on siDirect's own RNA/RNA scale: Freier 1986 nearest-neighbor parameters, helix initiation A = −10.8, CT = 100 µM, 100 mM Na⁺). Returns ranked candidates with sense/guide oligos (with UU 3' overhangs) and, per candidate, a ready shRNA cassette (sense–loop–antisense–Pol III terminator). Heuristic sequence rules only — no RNA-folding accessibility model and no transcriptome-wide off-target search. PREDICTED, NOT MEASURED. These are rule counts, not a regression, and no skill statistic is claimed for the ranking. Target-site accessibility is not modeled (no folding) and no transcriptome-wide off-target search is performed, so a top-ranked candidate is a starting point for a knockdown panel, not a predicted knockdown level. Valid for: mRNA targets. The seed Tm is computed on siDirect's own RNA/RNA scale (Freier 1986 parameters, helix initiation A = -10.8, CT = 100 µM, 100 mM Na+) and is not comparable with the DNA/DNA Tm reported elsewhere in the toolkit.Read-only
site_directed_mutagenesisDesign site-directed mutagenesis primers (QuikChange overlapping or Q5 back-to-back) for a base substitution, an amino-acid codon swap, or an insertion/deletion/delins. The edit can be given as fields or, more simply, by NAME in `mutation`: "E52K", "p.Glu52Lys", "c.155A>G", "c.76_78del", "c.76_77insGGA", "c.76_78dup". A named mutation is checked against the template — if the reference allele it states is not what is actually at that position, the call is refused and the real base or residue is quoted back, because a coordinate belonging to a different transcript or the other strand yields perfectly well-formed primers for the wrong base. `interpretedAs` in the response says which reading was designed.Read-only
solution_prepHow much to weigh out for a target molarity and volume: mass = C × V × MW. The molar mass comes from a named reagent's molecular FORMULA (computed, not transcribed) or from a value you supply. A bare name matching two hydration states is refused with both named rather than resolved to a guess — EDTA disodium dihydrate against the free acid is a 27% error that leaves a solution looking exactly like a solution.Read-only
synthesis_complexityMeasure the sequence features gene-synthesis vendors screen on — repeats (the single largest cause of synthesis failure), GC extremes, GC swings between adjacent windows, homopolymer runs and hairpin-forming inverted repeats — and report each against the threshold vendors publish. Returns measurements and named flags, never a success probability: refitted on the 303 REAL orders its authors publish (scripts/ssc-eval), the published classifier scores F1 0.866 against 0.878 for assuming every order succeeds, and transfers between ordering labs at AUC 0.423 — below chance. Pair with nonrepetitive_parts_design to fix the repeats it finds.Read-only
trace_diagnoseDiagnose a failing Sanger chromatogram: the longest usable window (Mott trimming), whether there is signal above the noise at all, whether more than one molecule is in the tube and from which base, whether quality collapsed at a homopolymer or tandem repeat the polymerase stuttered through, and whether a dye blob or the instrument's own separation is the problem rather than the DNA. Each finding carries the measurement it is based on, what that pattern is usually caused by, and what to do next. Reads .ab1, .abi and .scf.Read-only
trace_secondary_peaksRead a chromatogram for what the basecaller did not report: positions where a second dye is present under the called base (a heterozygote, or a contaminating template), and bases still visible in the scan past where base calling stopped. Both are arithmetic on the channel intensities already in the file — a ratio, and a peak position extrapolated from the median spacing — not a model. Returns the called sequence rewritten with IUPAC ambiguity codes, the extra bases, and every threshold that produced the answer. Takes the SCAN arrays, because the tail lies past the last peak location.Read-only
translateTranslate a nucleotide sequence to protein (single frame or all six frames; standard code).Read-only
variant_annotateOne-box variant lookup against MyVariant.info: accepts an rsID, chrom:pos:ref:alt (colon- or hyphen-separated), genomic HGVS ("chr17:g.7676154G>C"), or transcript HGVS c. ("NM_000546.6:c.215C>G" / "TP53:c.215C>G", bridged via the hgvs_convert tool). Returns a ClinVar significance summary, gnomAD exome/genome allele frequencies, and CADD/SIFT/PolyPhen2/REVEL pathogenicity predictor scores — each section explicitly null when that source has no data, never silently omitted. See the result's own "caveats" for real data-freshness limits (frozen gnomAD/CADD snapshots, periodic ClinVar snapshot).Read-only
variant_comparatorAlign a query to a reference and call variants (substitutions, insertions, deletions) in HGVS g. notation, with optional coding effects.Read-only
variant_to_constructTurn one variant into one buildable plan: verify the reference allele actually sits where the coordinate says, apply the edit, design site-directed mutagenesis primers to install it, design KASP/ARMS allele-specific primers to genotype it afterwards, and consolidate everything into a single oligo order table. Takes either a construct sequence with a 1-based position and ref/alt alleles (offline, deterministic), or an HGVS "c." description resolved through the MANE crosswalk and a live Ensembl exon map. A mismatched reference allele is refused with the bases that were actually found there, because a coordinate that is right for another isoform yields a perfectly valid primer set for the wrong base. Bases shared by both alleles are trimmed first, so a VCF-anchored pair is designed as the substitution or indel it actually is. Mutagenesis covers every class (a substitution, an insertion, a deletion and a multi-base replacement are all one interval replacement); KASP needs a single-base substitution's 3'-terminal base, so for an indel the genotyping half comes back as a named omission with the reason and the readout that does work, never as an empty list.Read-only
vector_library_getReturn one vector from the library: its GenBank accession and version, length, topology, organism/definition, complete sequence, and the full annotated feature table (type, label, 1-based inclusive start/end, strand, spliced length, and the location descriptor as the record wrote it). Accepts the library id, the vector name, or the accession. An unrecognized id is an error carrying the closest names — never an empty result.Read-only
vector_library_searchBrowse a curated library of publicly deposited, feature-annotated cloning and expression vectors — by name, category (E. coli cloning/expression, yeast, mammalian, plant binary, BAC/fosmid, recombineering, phage/M13), length window, or annotated feature (e.g. 'T7 promoter', 'ori', 'AmpR'). Each hit reports the vector's accession, length, topology and feature count; vector_library_get returns the sequence and the full feature table. A curated public-record set, NOT a vendor catalog — see the gate's notChecked.Read-only
verify_assemblyDeterministic self-check: given the same method/parts cloning_simulate would use (restriction-ligation, Gibson, Golden Gate, LIC, SLIC or In-Fusion/CPEC — optionally deriving a part by in-silico PCR first), re-derive the expected WHOLE product and diff it against a claimed final sequence. Returns pass/fail plus the exact position and nature of any discrepancy — not an opinion, the same deterministic simulation SeqBench already runs, run a second time as a check. A recipe that can give more than one molecule is checked against ALL of them and `matchedCandidate` names the one the claim matched: a non-directional ligation really does put the insert in both ways round (half the plate carries each), a vector cut more than twice offers more than one backbone, and a Gibson junction whose fragments already share terminal sequence has two honest readings (one homology arm, or a tandem repeat present twice). See verify_construct for a narrower, insert-only check that doesn't require declaring the vector/enzymes/method.Read-only
verify_constructRe-derive a construct's insert from the PCR (template + primers) claimed to have produced it, then check — independently of that claim — whether the expected insert actually appears (either orientation) in the claimed final construct, at what identity, and with exact mismatch positions if not. Optionally also checks for a premature stop in a declared reading frame. Primers may carry a non-templated 5' tail (a restriction site, a Gibson arm, a tag): a construct missing ONLY tail bases still passes, since that is exactly what digesting a tailed amplicon removes before ligation — see match.templateCoveragePct and match.unalignedIsTailOnly, and note the pass does not establish that the right enzyme made the cut. This re-derives from the claim's own stated inputs; it does not review the claim's prose.Read-only
virtual_gelPredict restriction-digest fragment sizes and their gel migration positions against a chosen DNA ladder.Read-only
volcano_plot_dataValidate a differential-expression table (gene, log2 fold-change, p-value/FDR) and compute -log10(p) plus up/down/non-significant counts at conventional default thresholds (|log2FC|>=1, p<=0.05), for the Volcano Plot visualization. Invalid rows (non-finite log2FC, or p-value outside (0,1]) are dropped and reported rather than failing the whole batch.Read-only
web_searchSearch the live web (via Tavily) for information not covered by SeqBench's own tools — recent literature, protocols, vendor/reagent info, general facts. Returns a short synthesized answer (if available) plus ranked source snippets with URLs. This does not run any bioinformatics calculation itself; use the dedicated tools for that.Read-only
whole_plasmid_verifyCompare a whole-plasmid sequencing CONSENSUS (the single circular FASTA a nanopore plasmid service returns) against the design it was supposed to be, and say whether they are the same molecule. Finds the rotation and strand itself — an assembler starts the circle wherever it landed and half the time returns the reverse complement, so the two almost never line up as written — then reports every difference in the DESIGN's own coordinates. Each difference carries the length of the identical-base run it sits in, because a 1 bp indel inside a homopolymer is both the commonest long-read basecalling artefact and what a real slippage mutation looks like, and a bare difference list cannot tell them apart. Pass the design as an annotated GenBank record and every difference is also read against its features: silent, missense with the codon number, or a premature stop. Complements sequencing_readback_verify, which aligns raw READS through minimap2; this one needs no reads and no sidecar.Read-only
workflowRun an ordered multi-tool pipeline over many records. One sequence value can be chained between steps.Read-only

Change history

No changes since the first observation. The first snapshot is the baseline.

Source listings
SourceListingFirst seenLast seenVersions
Official MCP Registrycom.seqbench/workbench2 Oct 20262 Oct 20261