$npx -y skills add parallel-web/parallel-agent-skills --skill parallel-data-enrichmentBulk data enrichment. Adds web-sourced fields (CEO names, funding, contact info) to lists of companies, people, or products. Use for enriching CSV files or inline data. Supports multi-turn: pass --previous-interaction-id from a prior research task to carry context forward.
| 1 | # Data Enrichment |
| 2 | |
| 3 | Enrich: $ARGUMENTS |
| 4 | |
| 5 | ## Before starting |
| 6 | |
| 7 | Inform the user that enrichment may take several minutes depending on the number of rows and fields requested. |
| 8 | |
| 9 | ## Optional: Suggest output columns |
| 10 | |
| 11 | If the user gave a vague intent ("enrich these companies with useful info") and you're not sure what columns to add, ask the API for a suggestion before kicking off the run: |
| 12 | |
| 13 | ```bash |
| 14 | parallel-cli enrich suggest "Find CEO and recent funding info" --json |
| 15 | ``` |
| 16 | |
| 17 | The response is an envelope: `{title, processor, enriched_columns, warnings}`. Extract just the **`enriched_columns` array** (not the whole envelope) and pass it as the value of `--enriched-columns` on `enrich run`, **in place of `--intent`** — the two flags are alternative ways to specify what to enrich, not combined. If `suggest` returned a `processor`, pass it through explicitly via `--processor` on the `run` call (it's a tuned recommendation for the schema). Skip this whole section if the user already specified the fields they want. |
| 18 | |
| 19 | > `enrich suggest` requires `parallel-cli` ≥ 0.3.0. If it errors with anything resembling `no such command` / `No such command` / `unknown command`, **do not bail** — skip the suggestion step, fall through to step 1 with `--intent`, complete the run, and mention `parallel-cli update` (or `pipx upgrade parallel-web-tools`) in the final response so the user picks up the feature next time. |
| 20 | |
| 21 | ## Step 1: Start the enrichment |
| 22 | |
| 23 | Use ONE of these command patterns (substitute user's actual data): |
| 24 | |
| 25 | For inline data: |
| 26 | |
| 27 | ```bash |
| 28 | parallel-cli enrich run --data '[{"company": "Google"}, {"company": "Microsoft"}]' --intent "CEO name and founding year" --target "output.csv" --no-wait --json |
| 29 | ``` |
| 30 | |
| 31 | For CSV file: |
| 32 | |
| 33 | ```bash |
| 34 | parallel-cli enrich run --source-type csv --source "input.csv" --target "output.csv" --source-columns '[{"name": "company", "description": "Company name"}]' --intent "CEO name and founding year" --no-wait --json |
| 35 | ``` |
| 36 | |
| 37 | If this is a **follow-up** to a previous research task and you have its `interaction_id`, add context chaining: |
| 38 | |
| 39 | ```bash |
| 40 | parallel-cli enrich run --data '...' --intent "..." --target "output.csv" --no-wait --json --previous-interaction-id "$INTERACTION_ID" |
| 41 | ``` |
| 42 | |
| 43 | The enrichment will run with the full context of that prior research — so you can enrich entities discovered earlier without restating what was already found. Note: enrichment does **not** itself produce a new `interaction_id`, so you cannot chain a further follow-up off of an enrichment. |
| 44 | |
| 45 | **IMPORTANT:** Always include `--no-wait` so the command returns immediately instead of blocking. |
| 46 | |
| 47 | Parse the `--json` output to extract `taskgroup_id` and `url`. The output is `{taskgroup_id, url, num_runs}` — there is no `interaction_id` field, do not look for one. Immediately tell the user: |
| 48 | |
| 49 | - Enrichment has been kicked off |
| 50 | - The monitoring URL where they can track progress |
| 51 | |
| 52 | Tell them they can background the polling step to continue working while it runs. |
| 53 | |
| 54 | ## Step 2: Poll for results |
| 55 | |
| 56 | Pick a concrete output path (e.g., `/tmp/enrichment-acme.json`). Note: the file is JSON regardless of the extension you choose — it's an array of `{input, output}` objects, not a CSV. Name it `.json` to avoid confusing yourself or the user. |
| 57 | |
| 58 | ```bash |
| 59 | parallel-cli enrich poll "$TASKGROUP_ID" --timeout 540 --output "/tmp/enrichment-<descriptive-name>.json" |
| 60 | ``` |
| 61 | |
| 62 | Important: |
| 63 | |
| 64 | - Use `--timeout 540` (9 minutes) to stay within tool execution limits |
| 65 | - The `--target` from step 1 is unused in `--no-wait` mode — only `--output` here determines where results are saved, and the file is always JSON |
| 66 | |
| 67 | ### If the poll times out |
| 68 | |
| 69 | Enrichment of large datasets can take longer than 9 minutes. If the poll exits without completing: |
| 70 | |
| 71 | 1. Tell the user the enrichment is still running server-side |
| 72 | 2. Re-run the same `parallel-cli enrich poll` command to continue waiting |
| 73 | |
| 74 | ## Response format |
| 75 | |
| 76 | **After step 1:** Share the monitoring URL (for tracking progress). |
| 77 | |
| 78 | **After step 2:** |
| 79 | |
| 80 | 1. Report number of rows enriched |
| 81 | 2. Preview first few rows from the output file (it's a JSON array of `{input, output}` objects) |
| 82 | 3. Tell the user the full path to the output file |
| 83 | |
| 84 | Do NOT re-share the monitoring URL after completion — the results are in the output file. |
| 85 | |
| 86 | ## Setup |
| 87 | |
| 88 | If `parallel-cli` is not found, install and authenticate: |
| 89 | |
| 90 | ```bash |
| 91 | /parallel:parallel-cli-setup |
| 92 | ``` |
| 93 | |
| 94 | If any `parallel-cli enrich` command returns `403`, tell the user balance is likely required. Offer to run `parallel-cli balance get`, and if needed a |