$npx -y skills add tt-a1i/matt-skills-with-to-goal --skill improve-codebase-architectureScan a codebase for deepening opportunities, present them as a visual HTML report, then grill through whichever one you pick.
| 1 | # Improve Codebase Architecture |
| 2 | |
| 3 | Surface architectural friction and propose **deepening opportunities** — refactors that turn shallow modules into deep ones. The aim is testability and AI-navigability. |
| 4 | |
| 5 | This command is _informed_ by the project's domain model and built on a shared design vocabulary: |
| 6 | |
| 7 | - Run the `/codebase-design` skill for the architecture vocabulary (**module**, **interface**, **depth**, **seam**, **adapter**, **leverage**, **locality**) and its principles (the deletion test, "the interface is the test surface", "one adapter = hypothetical seam, two = real"). Use these terms exactly in every suggestion — don't drift into "component," "service," "API," or "boundary." |
| 8 | - The domain language in `CONTEXT.md` gives names to good seams; ADRs in `docs/adr/` record decisions this command should not re-litigate. |
| 9 | |
| 10 | ## Process |
| 11 | |
| 12 | ### 1. Explore |
| 13 | |
| 14 | **Scope before you scan — YAGNI.** Deepening a module pays off by making future changes to it easier, so put extra weight on the parts of the codebase that have recently changed. Decide *where* to look before you look: |
| 15 | |
| 16 | - If the user named a direction — a module, a subsystem, a pain point — take it, and skip the inference below. |
| 17 | - Otherwise, walk back a good stretch of the commit history (`git log --oneline`) to find the codebase's hot spots — the files and areas that keep coming up — and let those paths pull your attention first. If the changes are scattered with no clear hot spot, widen the net. |
| 18 | |
| 19 | Read the project's domain glossary (`CONTEXT.md`) and any ADRs in the area you're touching first. |
| 20 | |
| 21 | Then use the Agent tool with `subagent_type=Explore` to walk the codebase. Don't follow rigid heuristics — explore organically and note where you experience friction: |
| 22 | |
| 23 | - Where does understanding one concept require bouncing between many small modules? |
| 24 | - Where are modules **shallow** — interface nearly as complex as the implementation? |
| 25 | - Where have pure functions been extracted just for testability, but the real bugs hide in how they're called (no **locality**)? |
| 26 | - Where do tightly-coupled modules leak across their seams? |
| 27 | - Which parts of the codebase are untested, or hard to test through their current interface? |
| 28 | |
| 29 | Apply the **deletion test** to anything you suspect is shallow: would deleting it concentrate complexity, or just move it? A "yes, concentrates" is the signal you want. |
| 30 | |
| 31 | ### 2. Present candidates as an HTML report |
| 32 | |
| 33 | Write a self-contained HTML file to the OS temp directory so nothing lands in the repo. Resolve the temp dir from `$TMPDIR`, falling back to `/tmp` (or `%TEMP%` on Windows), and write to `<tmpdir>/architecture-review-<timestamp>.html` so each run gets a fresh file. Open it for the user — `xdg-open <path>` on Linux, `open <path>` on macOS, `start <path>` on Windows — and tell them the absolute path. |
| 34 | |
| 35 | The report uses **Tailwind via CDN** for layout and styling, and **Mermaid via CDN** for diagrams where a graph/flow/sequence reliably communicates the structure. Mix Mermaid with hand-crafted CSS/SVG visuals — use Mermaid when relationships are graph-shaped (call graphs, dependencies, sequences), and hand-built divs/SVG when you want something more editorial (mass diagrams, cross-sections, collapse animations). Each candidate gets a **before/after visualisation**. Be visual. |
| 36 | |
| 37 | For each candidate, render a card with: |
| 38 | |
| 39 | - **Files** — which files/modules are involved |
| 40 | - **Problem** — why the current architecture is causing friction |
| 41 | - **Solution** — plain English description of what would change |
| 42 | - **Benefits** — explained in terms of locality and leverage, and how tests would improve |
| 43 | - **Before / After diagram** — side-by-side, custom-drawn, illustrating the shallowness and the deepening |
| 44 | - **Recommendation strength** — one of `Strong`, `Worth exploring`, `Speculative`, rendered as a badge |
| 45 | |
| 46 | End the report with a **Top recommendation** section: which candidate you'd tackle first and why. |
| 47 | |
| 48 | **Use CONTEXT.md vocabulary for the domain, and the `/codebase-design` vocabulary for the architecture.** If `CONTEXT.md` defines "Order," talk about "the Order intake module" — not "the FooBarHandler," and not "the Order service." |
| 49 | |
| 50 | **ADR conflicts**: if a candidate contradicts an existing ADR, only surface it when the friction is real enough to warrant revisiting the ADR. Mark it clearly in the card (e.g. a warning callout: _"contradicts ADR-0007 — but worth reopening because…"_). Don't list every theoretical refactor an ADR forbids. |
| 51 | |
| 52 | See [HTML-REPORT.md](HTML-REPORT.md) for the full HTML scaffold, diagram patterns, and styling guidance. |
| 53 | |
| 54 | Do NOT propose interfaces yet. After the file is written, ask the user: "Which of these would you like to explore?" |
| 55 | |
| 56 | ### 3. Grilling loop |
| 57 | |
| 58 | Once the user picks a candidate, run the `/grilling` skill to walk the decision tree with them — |