Genomic Intelligence

Hosted DNA language models: promoter, splice, enhancer, chromatin, expression, annotation

Community: Submitted by a user or imported; check the owner before granting accessOnlineAPI key requiredGlobalFreeRead-only

What it can do

    What data it sees

    Do you need an account

    An API key from the service settings is required

    Hosted DNA language models: promoter, splice, enhancer, chromatin, expression, annotation

    Server tool list (15)

    Raw names from tools/list. Only developers need these.

    list_modelsList available models for a task. Use to discover model ids before passing one as the `model` argument to a predict tool. The same catalog is also available as the resource `gi://models`. Returns a FLAT object — {task, default_model, models: [...]} — not the {data, meta} envelope the predict tools return. Each model carries a `bio_spec`, whose useful fields are `request_max_bp` (the enforced ceiling, 500,000 everywhere) and `context_window_bp` (what the model reads in one step — compare your sequence length against it: a shorter one is scored against a padded window). `trained_window_bp` is the fixed receptive field where there is no sliding window (9,198 for g0-expression). `request_max_bp` is the only one of the three that is a cap; the window fields describe what the model scores, not what the route accepts.
    fetch_ensembl_sequenceFetch a gene's reference sequence from Ensembl and store it. Returns a handle ({ref, name, length, preview, ...}). Pass the `ref` to predict_* tools — the bases stay server-side. For expression, use fetch_gene_for_expression instead (it prepares the TSS-centred window that model needs).
    fetch_regionFetch a genomic region by coordinates from Ensembl and store it. For "find the genes in chr8:127,680,000-127,800,000"-style requests: resolves a coordinate range to reference sequence and returns a handle ({ref, name, length, ...}) to pass to find_genes / predict_* — the bases stay server-side. Plus strand by default, which is what the gene-finder expects. For a gene by name use fetch_ensembl_sequence; for expression use fetch_gene_for_expression.
    fetch_gene_for_expressionFetch a gene's sequence prepared for expression prediction. Resolves the gene's TSS via Ensembl and returns the exact TSS-centred 9,198 bp window the expression model scores, as a handle to pass to predict_expression(sequence_ref=...). Because the window is exactly 9,198 bp, no `tss_index` is needed on that call.
    load_demo_sequenceLoad a bundled demo reference sequence and return a handle. The server ships one curated, task-correct positive control per task (list them via the gi://sequences resource) — e.g. `expression_hbb_k562` is a ready-to-use K562 expression window for predict_expression. Stores the demo and returns a handle to pass to a predict_* tool: no Ensembl fetch, no quota. Handy for smoke-testing a prediction end-to-end.
    store_inline_sequenceStore a human-pasted sequence and return a handle to re-use it. For a sequence you've already pasted into the conversation, this gives back a short handle so you can run several tasks on it without re-pasting the bases in each predict_* call. Note that the full sequence still passes through the LLM on THIS call — it does not save context on its own. For large sequences, prefer fetch_ensembl_sequence / fetch_gene_for_expression / load_local_fasta, which acquire the bases server-side and never round-trip them. A line-wrapped FASTA *body* may be pasted verbatim: whitespace is stripped before storing, so the handle's `length` counts bases and a later `tss_index` counts into the same string the API measures. (A FASTA `>` header line is not a sequence and is rejected by the API's alphabet check.)
    predict_promoterPredict promoter regions (G0). 300–500,000 bp. Returns the {data, meta} envelope: data.regions lists predicted promoters with start/end/score. 300 bp is the task floor for every promoter model. The default g0-promoter-2000bp scans a 2,000 bp context window, so a shorter (but ≥300 bp) sequence is still scored — against a window padded out to that size. Check the chosen model's bio_spec.context_window_bp via list_models to know whether it saw real sequence or padding.
    predict_splicePredict splice donor/acceptor sites (G0 BigBird). 100–500,000 bp. The model reads a 15,000 bp context window, so anything shorter is scored against a padded window — feed a whole transcript locus when you can. It is also strand-specific, and the wrong strand fails silently and plausibly — it returns sites at different positions, often still scoring above 0.9, not the near-zero scores once documented here. Nothing in the response flags it, so submit the transcript's own orientation (fetch_region takes `strand`).
    predict_enhancerPredict enhancer activity (G0 DeepSTARR). 50–500,000 bp. 50 bp is the task's admission floor (the API 422s below it), not a statement about what the model reads: enhancer models score a 249 bp context window, so 50–248 bp is accepted and scored against a padded window. For a meaningful call, submit at least the 249 bp context.
    predict_chromatinChromatin annotation across 919 features (G0 DeepSEA). 200–500,000 bp. The model reads a 1,000 bp context window; 200–999 bp is accepted and scored against a padded window.
    predict_expressionPredict a gene's expression from a TSS-centred window. Expression is cell-type-specific, so `description` (cell type / assay context, e.g. 'K562 cell line') is REQUIRED — the API rejects requests without it. The model scores exactly 9,198 bp centred on the TSS (±4,599). Two ways to supply that: - A sequence of exactly 9,198 bp already centred on the TSS. No `tss_index` needed — the midpoint is the only legal TSS. - A longer locus, 9,198–500,000 bp, plus `tss_index`: the 0-based offset of the TSS into it. The API cuts the window for you (sequence[tss_index-4599 : tss_index+4599]) and never scans for a TSS itself. Anything under 9,198 bp is rejected, here and by the API (422) — there is no padding or truncation fallback. `tss_index` is required for every other length, because a locus with no offset is indistinguishable from a mis-centred window. An offset that is merely WRONG (e.g. counted over a wrapped FASTA's characters, or against a chromosome coordinate instead of an offset into THIS sequence) still succeeds and scores the wrong window — verify meta.task_specific_counts.scored_window in the response. Easiest paths: fetch_gene_for_expression(gene) returns a ready-centred handle, and find_genes_and_predict_expression takes a raw region and finds each TSS for you.
    find_genesFind genes (transcript intervals) in a genomic region (async, ~8-25s). Takes 1,000–500,000 bp. The floor is the strictest of the scanning tasks: gene finding needs a region, not a site. (Only expression's 9,198 bp is higher, and that is a fixed window rather than a minimum region size.) Gene-finding: detects transcript boundaries (TSS + PolyA) and returns one interval per predicted transcript — start/end, strand, a confidence score, and predicted TSS/PolyA positions (BED-style feature intervals, not free-text notes). Use this for "what genes are here", "find / locate genes", or "annotate this region". Each transcript also carries its type (mRNA/lnc_RNA) and internal exon/intron/CDS structure in `exons`/`introns`/`cds` arrays, plus a browser-ready GFF3 track in `data.formats.gff3`. To get each gene's *expression* from a raw region, use find_genes_and_predict_expression instead — expression needs a per-gene TSS window, so predict_expression cannot run on a whole region. Submits an async job internally. With wait=True (default), blocks and streams progress, then returns the result {data, meta} — it never returns a job_id on this path. (If a generous block ceiling is exceeded it returns a timeout error, not a job handle.) With wait=False (detached), returns {data: {job_id, status: 'submitted'}} immediately — poll it with get_job.
    find_genes_and_predict_expressionFind genes in a sequence, then predict each gene's expression (composite). Server-side chaining in ONE call: finds genes (transcript intervals, with their TSS) in the sequence, then predicts expression off each discovered TSS in the given experimental context. This is the right tool whenever you want expression for a raw region or sequence — e.g. "find the genes in chr8:… and predict their expression in K562". predict_expression scores ONE TSS window and needs you to know where that TSS is (either a pre-centred 9,198 bp window or a `tss_index`); this tool discovers every gene's TSS itself. It has no 9,198 bp floor and no tss_index; it starts with gene finding, so it takes 1,000–500,000 bp. Runs async internally at every size (the annotate stage is slow even for small inputs), so progress always streams. With wait=True (default), blocks and streams progress, then returns the result {data, meta} — it never returns a job_id on this path. With wait=False (detached), returns {data: {job_id, status: 'submitted'}} immediately — poll it with get_job. Because it ends in expression, `description` (cell type / assay context) is REQUIRED.
    get_jobPoll an async job once. Returns the {data, meta} result if complete, a progress envelope if still running, or an error envelope if it failed.
    list_jobsList the caller's recent async jobs (also available as gi://jobs/recent).