SystemIntelPilot

Enterprise AI operator evaluation MCP server.

От сообщества: Добавлен пользователем или импортирован; проверьте владельца перед подключениемРаботаетБез входаГлобальныйБесплатноМожет изменять данные

Что умеет

  • Get Pilot Status: Get pilot status overview — cohort size, observation count, date range, data quality, active interventions. Computed from raw observations. Data is from a 50-operator synthetic pilot
  • Get Operator Profile: Get operator profile — operator details, measurements (5 canonical metrics computed from raw token observations with values, percentiles, status), and benchmark availability. Ope
  • Get Cohort Distribution: Get cohort metric distribution — min, p10, p25, median, p75, p90, max, mean, std, and outliers for a given metric across the 50-operator cohort. Computed from raw observations

Какие данные видит

Нужен ли аккаунт

Не нужен: сервер работает без входа

Enterprise AI operator evaluation MCP server. 27 tools (22 read + 5 write) for measuring, benchmarking, diagnosing, and interv ening on how human operators use AI tools across 5 canonical metrics.

Список инструментов сервера (27)

Технические названия из tools/list. Нужны только разработчикам.

get_pilot_statusGet pilot status overview — cohort size, observation count, date range, data quality, active interventions. Computed from raw observations. Data is from a 50-operator synthetic pilot (labeled synthetic).
get_operator_profileGet operator profile — operator details, measurements (5 canonical metrics computed from raw token observations with values, percentiles, status), and benchmark availability. Operator IDs are pseudonymous (e.g., op_001). Data is synthetic.
get_cohort_distributionGet cohort metric distribution — min, p10, p25, median, p75, p90, max, mean, std, and outliers for a given metric across the 50-operator cohort. Computed from raw observations.
get_composite_scoreGet developmental composite score (0-100) for an operator. Computed from raw metrics normalized via reference percentiles. Labeled DEVELOPMENTAL, not PERSONNEL. Weighted: leverage 30%, yield 30%, token_snr 20%, construction 20%. Data is synthetic.
get_composite_score_summaryGet cohort composite score summary — count, min, max, median, mean, Q1, Q3. Computed from per-operator scores. No individual rankings exposed. Label is DEVELOPMENTAL.
get_diagnosticsGet operator diagnostics — pattern detections and diagnoses computed from divergence analysis. All diagnoses are HYPOTHESIS, never fact.
get_data_qualityGet data quality summary — completeness, coverage, validity across the cohort. Computed from raw observations.
find_usage_operation_divergenceFind operators with usage-operation divergence. Computes usage percentile from raw token totals and compares to yield percentile. Returns all 50 operators with divergence class (LOW_USAGE_HIGH_OPERATION, HIGH_USAGE_LOW_OPERATION, etc.).
get_workflow_fitGet workflow fit analysis — operator/workflow fit scores across workflow stages.
get_intervention_statusGet all interventions — 12 active interventions with operator IDs, catalog IDs, reason patterns, target metrics, start dates, followup periods, and synthetic outcomes.
list_pilot_optionsList available pilot options — 5 canonical metrics, 15 eval families, 13 benchmark classes, 5 intervention types.
validate_pilot_configurationValidate a pilot configuration before deployment. Returns valid status with warnings and errors.
compare_operator_to_referenceCompare an operator to a reference population. Returns benchmark selection, comparison group, and metric comparison. Computed from raw metrics and reference field.
get_executive_dashboardGet executive dashboard info — the dashboard is a self-contained HTML file generated by the CLI (enterprise export dashboard --output file.html).
verify_changeVerify a measured change after intervention — pre/post comparison. Results are ASSOCIATION, never CAUSATION.
create_pilot_configurationGenerate a pilot configuration from parameters.
assign_interventionAssign a targeted intervention to an operator. REQUIRES AUTHORIZATION. Contact pilots@mos2es.org for pilot access.
close_interventionClose an intervention with outcome notes. REQUIRES AUTHORIZATION.
create_experimentCreate an experiment configuration. REQUIRES AUTHORIZATION.
record_workflow_observationRecord a workflow fit observation. REQUIRES AUTHORIZATION.
attach_outcome_datasetAttach external outcome dataset for join analysis. Outcome joins are ASSOCIATION, never CAUSATION. REQUIRES AUTHORIZATION.
get_operator_system_decompositionTwo-way ANOVA-style decomposition partitioning metric variance into operator effect, system effect, and operator×system interaction. Computed from raw observations grouped by platform. Shows whether operator capability or system choice drives performance.
get_lineage_chainGet the full lineage chain for an operator: STATE_A → BI_ACTION → AAI_TRANSFORMATION → BI_REDIRECTION → AAI_EXTENSION → COMMITTED_STATE → OUTCOME. Built from raw lineage and outcome data.
get_lineage_summaryGet lineage summary across the cohort — total lineages, workflow breakdown, average micro-eval metrics, outcomes linked. Computed from raw lineage data.
get_outcome_correlationCorrelate micro-eval metrics with outcome quality scores and cycle times through lineage. Computed via Pearson r from raw lineage + outcome data. Results labeled ASSOCIATION with evidence grade OBSERVATIONAL, never CAUSATION.
get_org_topologyOrganization-level AI topology map — team-level metric distributions, median canonical metrics per team, capability concentration (Gini coefficient), platform adoption, single-point-of-failure detection, cross-team complementarity. Computed from raw measurements.
get_operator_similarityNearest-neighbor operator search using percentile-rank normalization and Euclidean distance across 5 canonical metrics. Computed from raw measurements. Returns comparable operators/cohorts, NOT personality matching.
SystemIntelPilot: подключить к Claude, ChatGPT, Cursor · Connectors.fun