Project Discovery — Onboarding Orchestrator

SkillDev tools

Onboard a project through four discovery phases: Constitution, Architecture, Infrastructure, and Specification. Produces PRD, SRS, domain glossary, infrastructure context, and backlog access, then hands business maps and the master test plan to `project-context`. Use for set up this project, onboard this repo, connect to project, discover architecture, or create PRD/SRS. Do NOT use for incremental context refresh (`project-context`), writing tests, TMS documentation, running suites, adapting KATA (`adapt-framework`), or technical OpenAPI sync (`bun run api:sync`).

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Project Discovery — Onboarding Orchestrator skill

What this skill tells your AI

The instructions your AI receives, as published by upex-galaxy/agentic-qa-boilerplate in .agents/skills/project-discovery/SKILL.md and read by ahel’s review.

Turn an unknown codebase into a testable project. Four phases, always in order, gated on completion of the previous one. The output is a set of context files the rest of the skills (shift-left-testing, sprint-testing, test-automation, test-documentation, regression-testing) rely on.

The discovery is conversational: you read the code, ask when ambiguous, confirm before writing files. Never fabricate -- if you cannot verify a claim from the source, mark it as a "Discovery Gap" and move on.

Grounding methodology: IQL (Integrated Quality Lifecycle) — QA is continuous from requirement to release, not a gate at the end. The full rationale and step breakdown live in docs/methodology/IQL-methodology.md (shared across all QA skills). This skill does not depend on reading it — only point the user there if they ask why the discovery is structured this way.


Compact Rules

  • DO: run the four phases in order (Constitution → Architecture → Infrastructure → Specification), each gated on the previous. Show the output paths and wait for an explicit "Phase N complete" before continuing — never auto-chain.
  • DO NOT: write anything into the target repo. Discovery is read-only on it; .context/ is the only write target, and modifying the boilerplate itself is adapt-framework.
  • DO NOT: invent business entities, flows, requirements, or Jira/Xray field IDs and status names. Anything not verifiable from the source goes in the ## Discovery Gaps section that every output must carry.
  • DO: describe what the system DOES, not what product wants it to do. Discovery is reverse-engineering; a "to-be" PRD/SRS is out of scope — point the user at their own product workflow.
  • DO: lock the target repo path(s) before Phase 1 and block on ambiguity. A repo that is not cloned locally cannot be discovered from a URL — ask for the clone first.
  • WHEN the layout is split sibling repos: run the Phase 1 sub-steps once per repo and merge into ONE project-config.md, never interleaved. WHEN it is a monorepo: Phase 1 once project-wide, Phases 2-3 per package.
  • DO NOT: generate business maps, the feature catalog, or the master test plan here — those are project-context modes, which own their diff and overwrite approval. Exact API types are bun run api:sync.
  • DO NOT: create per-ticket PBI content or copy the backlog. Phase 4 produces only the backlog access recipe; the committed README.md and templates/ under .context/PBI/ stay untouched.
  • DO NOT: paste credentials or a detected secret into any discovery doc. Reference the .env key or the file path only; a hardcoded-secret hit is recorded as a HIGH risk with its path.
  • WHEN Phase 2 or 3 settles a test-architecture decision that is architectural AND hard to reverse (runner, isolation/parallelization, fixture and test-data strategy, auth-in-tests, selector contract, CI sharding): record it as an append-only ADR under .context/ADR/, drafted Proposed for the human to accept.
  • DO NOT: mix a discovery session with adapt-framework, and do not use this skill for incremental map refreshes — the write boundaries differ.
  • DO NOT: skip Phase 1 or its domain glossary on a fresh start. Downstream skills read the glossary as a precondition for ATP authoring and TC naming.
  • WHEN both a DB schema/migrations and ORM models exist: prefer the schema or migrations. ORM definitions drift from the live schema.
  • DO: mention the IQL methodology only if the user asks why the discovery is structured this way — never lecture someone who just wants the artifact.

Read full SKILL.md when: running any phase's sub-steps, applying a completion gate's content checks, or resolving the pre-adapt-framework prerequisite list.


Inputs

Canonical reading order when starting cold on a discovery run. Read in order; stop earlier when the scope is small enough that later inputs add no signal.

  1. Target project repo — path resolved at session start (see "Before starting: target repo location" below). Read code and any in-repo PRD. This is the primary source of truth — discovery is reverse-engineering, never aspirational design.
  2. Target repo's README.md and existing onboarding docs — fastest path to project intent, stack signals, and run commands before deep code reads.
  3. .context/ directory (if partial state exists from a prior discovery run) — informs Phase 0 resume decisions and prevents redundant work. Diff against current code before overwriting.
  4. .agents/project.yaml and .env.example — variable resolution patterns ({{PROJECT_KEY}}, env URLs, MCP names) that every downstream context file references.
  5. kata-manifest.json — registry of existing KATA Components + ATCs. Anchors what test surface the boilerplate already expects so discovery records gaps coherently.
  6. .agents/skills/agentic-qa-core/references/skill-composition-strategy.md — workflow context for downstream handoffs (project-context, adapt-framework, sprint-testing, test-documentation).
  7. Business / domain docs supplied by the user (Confluence, Notion exports, internal wikis) — secondary source for business model and glossary when in-repo signal is thin.

Subagent Dispatch Strategy

Orchestration & Session contracts: this skill follows agentic-qa-core/references/orchestration-doctrine.md (mandatory subagent dispatch — main thread is command center) AND agentic-qa-core/references/session-management.md (Phase 0 resume check, plan-first persistence at .session/<skill-slug>/<scope>/, archive on completion). Phase 0 (resume check) and Phase 1 (plan write) are NOT optional.

This skill is project-scope: no <scope> segment. Session state lives directly at .session/project-discovery/{plan.md, progress.md} per agentic-qa-core/references/session-management.md §3 + §9. This is the longest skill in the QA repo (1.5–4 hours, 4 hard-gate phases) and benefits most from per-phase checkpoints: if interrupted between Phase 2 (PRD/SRS) and Phase 3 (Infrastructure), resume reads progress.md and skips back to the first incomplete phase without re-prompting the user for already-confirmed scope.

This skill is compliant with the doctrine in AGENTS.md §"Orchestration Mode (Subagent Strategy)" and the session contract in .agents/skills/agentic-qa-core/references/session-management.md. Per-phase dispatch decisions live in Pick the scope first below: Fresh = heavy subagent delegation per phase; Boilerplate adoption = medium; Brownfield + Context refresh = main session only.


Phase 0 — Session resume check (MANDATORY, inline)

Before scope selection or any target-repo discovery, run the resume contract from agentic-qa-core/references/session-management.md §4:

  1. Check .session/project-discovery/progress.md.
  2. If it does NOT exist → proceed to "Before starting: target repo location" below, then "Pick the scope first" (which writes plan.md).
  3. If it DOES exist:
    • Read plan.md (chosen scope, target repo path, phase plan).
    • Read tail of progress.md (last completed phase + next planned phase).
    • Surface to the user: scope chosen, target repo, last completed phase, next phase, any open Discovery Gaps from the last entry.
    • Offer resume / restart / abort. On restart, archive to .session/.archive/<YYYY-MM-DD>-project-discovery-aborted/ before proceeding.

Resume is high-value here: Fresh onboarding (1.5–4h) crossing a session boundary without resume re-runs Phase 1 from scratch, re-prompting target paths the user already confirmed.


Before starting: target repo location

/project-discovery runs read-only against a project under test — the target repo — that is NOT this boilerplate. Before Phase 1 starts, lock down where the target lives. Block Phase 1 if the target path is ambiguous.

LayoutWhat to declareHow to detect
Monorepo (single repo contains FE + BE)Absolute or relative path from this repoCheck the candidate path for pnpm-workspace.yaml, turbo.json, nx.json, lerna.json, or a top-level package.json with no deps of its own
Split sibling repos (FE and BE cloned separately)One path per repo (or a common parent dir)Look at ../-level siblings with plausible names (*-backend, *-frontend, *-api, *-web); confirm with the user
Remote (not cloned yet)Repo URL + branch, then ask the user to clone locally before Phase 1gh repo view only returns metadata; real discovery needs local file access — do not try to discover from a URL

Record the resolved path(s) in .context/project-config.md §Repositories during Phase 1 sub-step 1 (Project Connection). Every <target-repo> reference in later phases resolves to the path declared here.

If the layout is "split sibling repos", run Phase 1 sub-steps once per repo and merge findings into a single project-config.md; do not interleave.


Pick the scope first

All projects go through the same 4 phases, but depth varies. Pick once, then follow the common pipeline.

ScenarioInputPhases to runTypical depthContext weight & subagent hint
Fresh onboarding (greenfield or unseen project)Repo URL or local path(s), no existing context files1 -> 2 -> 3 -> 4, then project-context refresh-allFull discovery. Business maps and test strategy are generated by their dedicated skill. After context completion, run adapt-framework.Heavy. Delegate each phase's code survey to a dedicated subagent.
Boilerplate adoption (this repo adopted for a new project)Target app repo(s), this repo as the test framework1 (project-connection) -> 3, then project-context for missing mapsSkip Phase 2 or 4 only when their required artifacts already exist. Verify files on disk before adapt-framework.Medium. Delegate Phase 1 and Phase 3 per package for monorepos.
Brownfield (project already documented, tests missing)Existing .context/ partially filled2 (gaps) -> 3 (gaps) -> 4 (gaps), then project-context for stale mapsFill discovery gaps here; refresh map artifacts in their owning skill.Light. Main session unless gaps span many files.
Context refreshUser asks to regenerate a business map or master test planRedirect to the matching project-context modeThis skill does not refresh those artifacts. For PBI access changes, re-run Phase 4. For exact OpenAPI types, use bun run api:sync.Minimal. Handoff only.

Default to "Fresh onboarding" when in doubt. Confirm the scope with the user before starting Phase 1.

After scope confirmation, write .session/project-discovery/plan.md per agentic-qa-core/references/session-management.md §6. The phase breakdown ends at Phase 4; record project-context refresh-all as the post-discovery handoff, not as a discovery phase.


Workflow — the 4-phase pipeline

Phase 1: Constitution        -> Phase 2: Architecture       -> Phase 3: Infrastructure    -> Phase 4: Specification
(who/what/why)                 (PRD + SRS)                    (backend/frontend/infra)       (PBI mapping)
                |                      |                              |                              |
   .context/business/            .context/PRD/*.md           .context/infrastructure/*.md     .context/PBI/ACCESS.md
   business-model.md            .context/SRS/*.md
   domain-glossary.md
   project-config.md

                                                 |
                                                 v
                                    project-context (separate skill)
                                    data -> features -> api -> test-plan
                                    `bun run api:sync` remains the technical
                                    OpenAPI type pipeline.

KATA adaptation is a separate skill: adapt-framework. It runs after discovery and context outputs exist.

Each phase has a completion gate: before moving on, the required output files must exist on disk with non-placeholder content. Ask the user to confirm after each phase; never auto-chain.

Phase 1 — Constitution (who, what, why)

Goal: make the project legible. Outputs are read by every future session.

Four sub-steps, in order:

  1. Project Connection -- repo paths, tech stack detection, environment URLs, credentials from .env, team contacts.
  2. Project Assessment -- current testing maturity (frameworks in place, CI presence, lint/typecheck, coverage). Produces a risk profile.
  3. Business Model Discovery -- problem statement, target users, value proposition, revenue model (if any). Business Model Canvas recommended.
  4. Domain Glossary -- core entities, relationships, state machines, enumerations, UI-label vs code-identifier mapping.

Completion gate: .context/business/business-model.md, .context/business/domain-glossary.md, .context/project-config.md all exist and are non-empty. Plus a ## Project Assessment (Phase 1) block in canonical AGENTS.md. Sanity-check content — these are soft gates, surfaced to the human as warnings, not hard aborts:

  • domain-glossary.md contains at least 5 core-entity subsections (grep ^### yields 5+ matches, ignoring top-level H3s from "Enumerations" etc. — aim for real entities).
  • business-model.md cites at least one concrete source (Source: or Found in: literal appears 3+ times).
  • project-config.md has a ## Tech Stack section AND a ## Environments section.

After the automated sanity check, show the human the output paths and wait for explicit "Phase 1 complete, continue" before moving on.

Read references/phase-1-constitution.md when running any Phase 1 sub-step. Contains the discovery process, stack-detection commands, required output sections, and quality checklists.

Phase 2 — Architecture (PRD + SRS)

Goal: produce the Product and Software Requirements docs from code (not the other way round -- that is the "creation" direction, this is the "discovery" direction).

PRD sub-steps (run first, in parallel or sequentially — user choice):

  1. Executive Summary -- problem, solution, success metrics, scope.
  2. User Personas -- roles, permissions, primary/secondary users, role hierarchy.
  3. User Journeys -- critical paths through the UI, route map, journey diagrams.

Feature catalog is post-discovery. project-context mode features owns .context/business/business-feature-map.md. Do not generate it here.

SRS sub-steps (run after PRD, serially):

  1. Architecture Specs -- C4 context and container diagrams, component structure, database schema, external services, security model.
  2. Functional Specs -- FR-N entries with preconditions, business rules, validations, state machines.
  3. Non-Functional Specs -- performance budgets, security posture, reliability (RTO/RPO), scalability, observability, compliance.

API contracts are NOT an SRS output. The technical surface is owned by bun run api:sync; the business angle is owned by project-context mode api. Phase 2 records only the spec location or a Discovery Gap.

Test-architecture ADR seeding (Phase 2 SRS + Phase 3). When the Architecture Specs / Infrastructure sub-steps settle a hard-to-reverse test-architecture decision — test runner/framework, isolation & parallelization model, fixture/test-data strategy, auth-in-tests, selector/data-testid contract, exploratory-vs-scripted boundary, CI sharding — promote each one that passes the two-gate test (architectural AND hard to reverse) to a standalone ADR-NNNN-<slug>.md in .context/ADR/, and reference it from architecture.md / infrastructure/. Greenfield: you are ENCODING the decision; brownfield: you are RECORDING the one you discovered. Follow agentic-qa-core/references/adr-doctrine.md (detection + authoring) and .context/ADR/README.md (template + lifecycle). AI drafts Proposed; the human accepts.

Completion gate: .context/PRD/executive-summary.md, user-personas.md, user-journeys.md, .context/SRS/architecture.md, functional-specs.md, non-functional-specs.md all exist. API contract source is recorded in .context/project-config.md. business-feature-map.md remains a post-discovery project-context output. Soft content checks:

  • architecture.md contains at least one ```mermaid block AND one of (## Data Flow, ## Database Schema, ## Component Structure).
  • functional-specs.md contains at least one FR- identifier and one BR- identifier.
  • user-personas.md lists at least 2 role entries (### or table rows with role names).

Show outputs to the human and wait for "Phase 2 complete, continue" before moving on.

Read references/phase-2-prd.md when working on any PRD doc. Read references/phase-2-srs.md when working on any SRS doc. They are independent -- do not load both unless you are straddling both sides.

Phase 3 — Infrastructure

Goal: make the project runnable and deployable for the test environment.

Three sub-steps:

  1. Backend Discovery -- language, framework, database, ORM, auth, dependency manager, run/test commands, migrations, env vars.
  2. Frontend Discovery -- framework, bundler, routing, state management, design system, component library, test IDs strategy.
  3. Infrastructure Mapping -- CI/CD providers, deployment targets, environments (dev/staging/prod), infra-as-code, monitoring, rollback procedure.

Completion gate: .context/infrastructure/backend.md, frontend.md, infrastructure.md all exist with the key facts (auth flow, test commands, deploy URLs) filled in. Soft content checks:

  • backend.md AND frontend.md each contain a ## Runtime (or ## Build Configuration) section AND a commands block (bash fenced) covering install + run.
  • infrastructure.md lists environments explicitly (| Staging | or | Production | table row).
  • At least one auth-flow pointer exists in backend.md (e.g., mentions /auth/login, session, JWT, cookie, OAuth).

Show outputs to the human and wait for "Phase 3 complete, continue" before moving on.

Read references/phase-3-infrastructure.md when running any Phase 3 sub-step. Contains framework-detection heuristics, required sections per artifact, and common gotchas (SSR vs CSR, edge vs serverless, monorepo vs split repos).

Phase 4 — Specification (Backlog mapping)

Goal: hook the testing framework into the team's issue tracker without duplicating content.

One sub-step:

  1. PBI Backlog Mapping -- connect to {{ISSUE_TRACKER}} via [ISSUE_TRACKER_TOOL], discover project key, map hierarchy (Epic/Story/Task/Bug), record queries used to fetch tickets. Output: .context/PBI/ACCESS.md (backlog access recipe). NEVER write .context/PBI/README.md — it is a committed framework document (tier doctrine + gitignore ladder), not a discovery output; same for the committed templates/ skeletons.

Per-ticket PBI is NOT generated by this skill. It is materialized later by /sprint-testing via bun run jira:sync-issues get <KEY> --include-comments, which writes the canonical synced tree .context/PBI/epics/EPIC-<KEY>-<slug>/stories/STORY-<KEY>-<slug>/ (Module = Epic, 1:1). Those local .md files are a READ-ONLY cache of Jira (Jira = source of truth). This skill does NOT create per-ticket story.md — it only sets up the backlog access recipe (ACCESS.md).

Completion gate: .context/PBI/ACCESS.md exists with project key + auth recipe. Soft content checks:

  • PBI/ACCESS.md contains the configured {{PROJECT_KEY}} literal AND a ## Common Queries section (or JQL / WIQL snippet).
  • .context/PBI/README.md and .context/PBI/templates/ untouched (framework-owned, committed).

Show outputs to the human and wait for "Phase 4 complete" before emitting the project-context handoff.

Read references/phase-4-specification.md when running Phase 4. Contains issue-tracker connection recipes, query conventions, and the ACCESS.md structure.

Business-context handoff

Business maps and the master test plan are not generated here. After Phase 4, open a clean session and invoke project-context mode refresh-all. It owns the deterministic sequence data -> features -> api -> test-plan, including every CREATE/UPDATE approval gate. Exact OpenAPI types remain owned by bun run api:sync.

After those outputs exist, invoke adapt-framework to wire this boilerplate to the target stack.


Per-phase progress + Archive

After each phase passes its completion gate AND the user confirms "Phase N complete", append a phase entry to .session/project-discovery/progress.md. Entries end at Phase 4; the next action is the separate project-context skill.

After Phase 4 passes, archive the project-discovery session per agentic-qa-core/references/session-management.md §8 and record the project-context refresh-all handoff. Context generation has its own lifecycle and does not keep this session open.

On Phase-gate REJECT (user marks a phase incomplete or finds a Discovery Gap that blocks), archive does NOT run. The working directory stays so resume picks up at the failing gate.


Next recommended steps (emit after Phase 4 completes)

Discovery populates PRD, SRS, glossary, infrastructure, and backlog access. It does not invoke project-context, which is token-heavy and best run in a clean session.

When Phase 4 is confirmed complete, print this block to the user verbatim:

Discovery complete. `/project-discovery` has populated:
- .context/business/business-model.md, domain-glossary.md
- .context/project-config.md
- .context/PRD/executive-summary.md, user-personas.md, user-journeys.md
- .context/SRS/architecture.md, functional-specs.md, non-functional-specs.md
- .context/infrastructure/backend.md, frontend.md, infrastructure.md
- .context/PBI/ACCESS.md

**Recommended next skill** (run in a clean session):

`project-context` mode `refresh-all`

It runs data -> features -> api -> test-plan in dependency order and can be re-run whenever project context becomes stale.

After it completes, invoke `adapt-framework` to wire KATA against the target stack.

Do not auto-chain the handoff inside this session. Context generation needs its own token budget and approval lifecycle.

Pre-adapt-framework checklist

Before the user invokes adapt-framework, verify every file below is on disk. Missing business maps route to the matching project-context mode.

  • .context/PRD/ populated (at least README.md) AND .context/business/business-model.md or domain-glossary.md present
  • .context/SRS/architecture.md
  • .context/infrastructure/backend.md and .context/infrastructure/frontend.md
  • .context/business/business-data-map.md
  • API contract source: one of api/openapi-types.ts (non-stub) OR reachable OpenAPI spec URL OR .context/business/business-api-map.md (business-angle fallback)
  • .env.example (and .env either present or created during adapt-framework)

Handoff line to print to the user:

Discovery handoff complete. Run project-context refresh-all, then invoke adapt-framework when the six prerequisites are present.


Stack-specific discovery rules

Base stack detection (package.json → Node, pyproject.toml → Python, go.mod → Go, next.config.* → Next.js, etc.) is a baseline skill any AI has. This section only lists actions the skill should take based on what is detected — rules that are not obvious from general programming knowledge.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
22
Forks
13
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
project-discovery-upex-galaxy
Source
github.com/upex-galaxy/agentic-qa-boilerplate