Mozilla Data Collective

Search the Mozilla Data Collective catalog of ethically sourced AI training datasets.

Community: Submitted by a user or imported; check the owner before granting accessDegradedNo sign-inGlobalFreeRead-only

What it can do

    What data it sees

    Do you need an account

    No: the server works without sign-in

    Search the Mozilla Data Collective catalog of ethically sourced AI training datasets.

    Server tool list (3)

    Raw names from tools/list. Only developers need these.

    searchSearch the Mozilla Data Collective catalog of AI training datasets by natural-language query, optionally narrowed by task, language, license, format, price, sample availability or publish date. Returns matching datasets as {id, title, url}; pass an id to the fetch tool for full details. Call list_filters first if you intend to filter — filter values must match the catalog exactly.
    fetchFetch the full public details of one Mozilla Data Collective dataset by id or slug: description, organization, task, locale, license, format, size, pricing, and its page URL.
    list_filtersList every value the search tool's filters accept: the tasks, locales, licenses and formats present in the catalog, plus the sort and date-range options. Task and license values are abbreviations, so taskLabels and licenseLabels spell them out. Filter values are matched exactly, so call this before filtering a search rather than guessing values. Takes no arguments.
    Mozilla Data Collective: connect to Claude, ChatGPT, Cursor · Connectors.fun