---
name: orbit-ab-test
description: Run an A/B comparison of the same question with and without Orbit (GitLab Knowledge Graph), then produce a structured feedback report for the Orbit design partner program. Use when the user asks to A/B test a prompt, compare Orbit vs no Orbit, or generate Orbit feedback.
---

# Orbit A/B Test

You are running a controlled comparison for the Orbit design partner program. The user gives you one question (the prompt under test). You answer it twice under different conditions, measure both runs, and produce one structured report.

## Ground rules

- The two runs must be isolated. If your environment supports subagents or fresh sessions (for example the Agent tool in Claude Code), run each variant in its own subagent with an identical prompt. Never let Run B reuse anything Run A discovered.
- A prompt-level ban is not enough for Run B if the Orbit MCP server is configured, because the subagent can still see those tools. Before Run B, ask the user to disable the Orbit MCP server for the session, or run B in a session that never had it. Note in the report whether the control run had Orbit tools visible.
- If you cannot isolate runs, tell the user to invoke this skill twice in fresh sessions (once per variant) and then ask you to assemble the report from both transcripts.
- Do not soften the result. If Orbit's answer was worse, slower, or wrong, the report says so. Negative results are exactly what the program wants.

## Run A: with Orbit

Instruct the subagent to answer the user's question using Orbit as the primary source: the `glab` CLI (preferred) and/or the Orbit MCP tools for schema exploration and graph queries. File reading is allowed to verify findings. Record:

- Start and end time (wall clock)
- Number of tool calls, and how many were Orbit queries
- Approximate tokens consumed, if your environment reports usage
- The final answer

## Run B: without Orbit

Instruct the subagent to answer the identical question with Orbit forbidden: no Orbit MCP tools, no `glab` Orbit commands. It may use anything else it normally would (file reading, grep, repository search). Record the same measurements.

## Report

Output exactly this structure, in markdown:

```
## Orbit A/B Test Report

**Question under test:** <the prompt, verbatim>
**Agent + model:** <e.g. Claude Code / model name>
**Date:** <date>
**Repo/group scope:** <REDACT before sharing publicly>

| Measure | A: With Orbit | B: Without Orbit |
|---|---|---|
| Wall time | | |
| Tool calls (Orbit queries) | | n/a |
| Approx. tokens | | |
| Files/entities found | | |

### Answer A (with Orbit)
<summary of the answer, key findings>

### Answer B (without Orbit)
<summary of the answer, key findings>

### Agent assessment (1-5 each, one line of justification)
- Completeness: A __ / B __: <why>
- Groundedness (real files, no guesses): A __ / B __: <why>
- Effort required: A __ / B __: <why>

### What A found that B missed (or vice versa)
<concrete list>

### Human verdict (partner fills in)
- Which answer would you act on, and why:
- Anything Orbit got wrong or missed:
```

## After the report

1. Remind the user: the agent's self-assessment is approximate. The **human verdict** line is the one the Orbit team weighs most.
2. Remind the user to **redact private file paths, symbol names, and code** before posting anywhere public.
3. Tell them where it goes: paste the redacted report into the Orbit feedback epic at https://gitlab.com/groups/gitlab-org/orbit/-/work_items/4, or bring it to the weekly program sync if it can't be shared publicly.
