paper-mcp

Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.

Community: Submitted by a user or imported; check the owner before granting accessOnlineNo sign-inGlobalFreeRead-only

What it can do

    What data it sees

    Do you need an account

    No: the server works without sign-in

    Search arXiv/Semantic Scholar/OpenAlex + medical evidence (PubMed/Europe PMC) + LaTeX/PDF tools.

    Server tool list (41)

    Raw names from tools/list. Only developers need these.

    search_papersSearch academic papers. Returns normalized hits with a short abstract preview; call get_paper for the full record.
    search_allAggregated search across arXiv, Semantic Scholar and OpenAlex at once. Fans out concurrently, de-duplicates the same work across corpora (by DOI or title) and re-ranks with Reciprocal Rank Fusion, so papers found by several sources rank highest. Each hit lists which `sources` found it and an `ids` map ({source: id}) you can pass to get_paper / read_paper / the citation tools. Prefer this over search_papers for a broad lookup.
    search_medicalEvidence-graded MEDICAL literature search (PubMed + Europe PMC). Unlike search_all (generic, ranks high-cited reviews/guidelines above trials), this filters by research type via PubMed Publication-Type tags and re-ranks by the evidence pyramid (meta-analysis / systematic review > RCT > cohort > ...), so the actual clinical trials surface first. Open-access full text is pulled from Europe PMC by PMID. `query` should be English keyword/boolean text (PubMed maps it); do natural-language/multilingual understanding upstream. Returns hits with pmid/doi/study_type/evidence_level/citations/abstract and, when open-access, fulltext.
    get_paperFetch one paper by id, with full abstract and PDF link.
    search_by_authorFind papers by a specific author, newest first.
    list_recentList the latest papers in a subject category, newest first.
    list_categoriesList common subject category codes for filtering/recent.
    read_paperRead a paper's full text. format='markdown' (default, body with formulas as $LaTeX$), 'html' (raw LaTeXML HTML), or 'latex' (the original LaTeX manuscript from the e-print source). arXiv only; id like 2401.01234.
    list_paper_sourcesList available paper corpora.
    recognize_formulaRecognize a math formula from an image and return LaTeX. Provide image_url (downloaded server-side) OR image_base64. model: deepseek-ocr (default), paddleocr-vl, or texify. Returns {latex, model, elapsed_ms}.
    recognize_tableRecognize a table from an image and return LaTeX tabular code. Provide image_url OR image_base64. model: deepseek-ocr (default), paddleocr-vl, or texify. Returns {latex, model, elapsed_ms}.
    list_ocr_modelsList the OCR models available for recognize_formula / recognize_table.
    lint_latexLint a LaTeX snippet: report errors and return an auto-fixed version. Input `code` (the LaTeX source). Returns {errors, fixed_code, summary_en, summary_zh, elapsed_ms}.
    extract_pdfExtract a PDF to clean Markdown/LaTeX text via MinerU (great for papers behind no open-access full text — give the user's PDF and get readable text back). Provide pdf_url (downloaded server-side, SSRF-guarded) OR pdf_base64. formula/table toggle math/table reconstruction. Returns {task_id, status, cached, content, chars}: a recently-seen (cached) or small PDF comes back with `content` in one call; a fresh PDF (MinerU is GPU-heavy, minutes) returns status='running' + a task_id — then call extract_pdf_result(task_id) to fetch the text.
    extract_pdf_resultFetch the result of an extract_pdf job by task_id. Returns {task_id, status, content, chars}: `content` is the extracted text once status='done'; while still 'running' content is null — call again shortly. Results expire server-side, so fetch reasonably soon.
    get_paper_citationsSemantic Scholar: papers that CITE this one (forward citation graph). id accepts S2 id / DOI: / ARXIV: / CorpusId:.
    get_paper_referencesSemantic Scholar: papers this one REFERENCES (its bibliography). id accepts S2 id / DOI: / ARXIV: / CorpusId:.
    get_paper_authorsSemantic Scholar: the authors of a paper (with h-index, paper/citation counts).
    match_paper_titleSemantic Scholar: find the single paper whose title best matches the given text (exact-match lookup).
    autocomplete_papersSemantic Scholar: autocomplete paper titles for a partial query (fast type-ahead).
    search_papers_bulkSemantic Scholar: bulk paper search (up to 1000 hits, sortable e.g. 'citationCount:desc' or 'publicationDate:desc', with a continuation token). Filters: fields_of_study, year (e.g. '2020-2024'), venue, publication_types, open_access_pdf.
    get_papers_batchSemantic Scholar: fetch many papers at once by id (S2/DOI:/ARXIV:/CorpusId:), up to ~500 per call.
    search_authorsSemantic Scholar: search for authors by name; returns profiles with h-index and paper/citation counts.
    get_authorSemantic Scholar: a single author's profile by id.
    get_author_papersSemantic Scholar: all papers by a given author id, newest first.
    get_authors_batchSemantic Scholar: fetch many authors at once by id.
    search_snippetsSemantic Scholar: search INSIDE paper full text and return matching text snippets (not just titles/abstracts).
    recommend_papers_for_paperSemantic Scholar: recommend papers similar to one paper. pool='recent' (last open corpus) or 'all-cs' (all of CS). If the 'recent' pool yields nothing (common for older papers), it automatically retries the 'all-cs' pool.
    recommend_papers_from_examplesSemantic Scholar: recommend papers from positive (and optional negative) example paper ids.
    list_dataset_releasesSemantic Scholar Datasets: list all available release ids (dated snapshots of the full corpus).
    get_dataset_releaseSemantic Scholar Datasets: which datasets a release contains (papers, abstracts, citations, embeddings, s2orc, tldrs…). release_id defaults to 'latest'.
    get_dataset_download_linksSemantic Scholar Datasets: get download links (presigned URLs) for one dataset in a release. Needs the API key.
    get_dataset_diffsSemantic Scholar Datasets: incremental diff (added/updated/deleted) for a dataset between two releases. Needs the key.
    get_openalex_workOpenAlex: fetch one work's full record (316M-work, all-field corpus). id accepts OpenAlex Wxxxx, a DOI, or an arXiv id.
    get_openalex_citationsOpenAlex: papers that CITE this work (forward citation graph), most-cited first.
    get_openalex_referencesOpenAlex: the works this one REFERENCES (its bibliography).
    search_openalex_authorsOpenAlex: search authors; returns profiles with h-index, i10-index, works/citation counts and institutions.
    search_openalex_institutionsOpenAlex: search institutions (universities, labs) with ROR id, country, works/citation counts.
    search_openalex_worksOpenAlex: advanced filtered work search. Filters: from_year, to_year, is_oa (open access only), min_citations, institution_id. sort_by: relevance|newest|cited.
    get_openalex_trendsOpenAlex: publication-trend analytics for a query — counts grouped by year (default), or by 'institutions.id', 'authorships.author.id', 'open_access.is_oa', 'type', 'language'. Returns aggregate counts only (cheap, no rows).
    list_openalex_topicsOpenAlex: search the topic taxonomy (~4500 topics) to find the right subject term for filtering or recent-work queries.
    paper-mcp: connect to Claude, ChatGPT, Cursor · Connectors.fun