$npx -y skills add mjunaidca/mjs-agent-skills --skill docxComprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction. When Claude needs to work with professional documents (.docx files) for: (1) Creating new documents, (2) Modifying or editing content,
| 1 | # DOCX creation, editing, and analysis |
| 2 | |
| 3 | ## Overview |
| 4 | |
| 5 | A user may ask you to create, edit, or analyze the contents of a .docx file. A .docx file is essentially a ZIP archive containing XML files and other resources that you can read or edit. You have different tools and workflows available for different tasks. |
| 6 | |
| 7 | ## Workflow Decision Tree |
| 8 | |
| 9 | ### Reading/Analyzing Content |
| 10 | Use "Text extraction" or "Raw XML access" sections below |
| 11 | |
| 12 | ### Creating New Document |
| 13 | Use "Creating a new Word document" workflow |
| 14 | |
| 15 | ### Editing Existing Document |
| 16 | - **Your own document + simple changes** |
| 17 | Use "Basic OOXML editing" workflow |
| 18 | |
| 19 | - **Someone else's document** |
| 20 | Use **"Redlining workflow"** (recommended default) |
| 21 | |
| 22 | - **Legal, academic, business, or government docs** |
| 23 | Use **"Redlining workflow"** (required) |
| 24 | |
| 25 | ## Reading and analyzing content |
| 26 | |
| 27 | ### Text extraction |
| 28 | If you just need to read the text contents of a document, you should convert the document to markdown using pandoc. Pandoc provides excellent support for preserving document structure and can show tracked changes: |
| 29 | |
| 30 | ```bash |
| 31 | # Convert document to markdown with tracked changes |
| 32 | pandoc --track-changes=all path-to-file.docx -o output.md |
| 33 | # Options: --track-changes=accept/reject/all |
| 34 | ``` |
| 35 | |
| 36 | ### Raw XML access |
| 37 | You need raw XML access for: comments, complex formatting, document structure, embedded media, and metadata. For any of these features, you'll need to unpack a document and read its raw XML contents. |
| 38 | |
| 39 | #### Unpacking a file |
| 40 | `python ooxml/scripts/unpack.py <office_file> <output_directory>` |
| 41 | |
| 42 | #### Key file structures |
| 43 | * `word/document.xml` - Main document contents |
| 44 | * `word/comments.xml` - Comments referenced in document.xml |
| 45 | * `word/media/` - Embedded images and media files |
| 46 | * Tracked changes use `<w:ins>` (insertions) and `<w:del>` (deletions) tags |
| 47 | |
| 48 | ## Creating a new Word document |
| 49 | |
| 50 | When creating a new Word document from scratch, use **docx-js**, which allows you to create Word documents using JavaScript/TypeScript. |
| 51 | |
| 52 | ### Workflow |
| 53 | 1. **MANDATORY - READ ENTIRE FILE**: Read [`docx-js.md`](docx-js.md) (~500 lines) completely from start to finish. **NEVER set any range limits when reading this file.** Read the full file content for detailed syntax, critical formatting rules, and best practices before proceeding with document creation. |
| 54 | 2. Create a JavaScript/TypeScript file using Document, Paragraph, TextRun components (You can assume all dependencies are installed, but if not, refer to the dependencies section below) |
| 55 | 3. Export as .docx using Packer.toBuffer() |
| 56 | |
| 57 | ## Editing an existing Word document |
| 58 | |
| 59 | When editing an existing Word document, use the **Document library** (a Python library for OOXML manipulation). The library automatically handles infrastructure setup and provides methods for document manipulation. For complex scenarios, you can access the underlying DOM directly through the library. |
| 60 | |
| 61 | ### Workflow |
| 62 | 1. **MANDATORY - READ ENTIRE FILE**: Read [`ooxml.md`](ooxml.md) (~600 lines) completely from start to finish. **NEVER set any range limits when reading this file.** Read the full file content for the Document library API and XML patterns for directly editing document files. |
| 63 | 2. Unpack the document: `python ooxml/scripts/unpack.py <office_file> <output_directory>` |
| 64 | 3. Create and run a Python script using the Document library (see "Document Library" section in ooxml.md) |
| 65 | 4. Pack the final document: `python ooxml/scripts/pack.py <input_directory> <office_file>` |
| 66 | |
| 67 | The Document library provides both high-level methods for common operations and direct DOM access for complex scenarios. |
| 68 | |
| 69 | ## Redlining workflow for document review |
| 70 | |
| 71 | This workflow allows you to plan comprehensive tracked changes using markdown before implementing them in OOXML. **CRITICAL**: For complete tracked changes, you must implement ALL changes systematically. |
| 72 | |
| 73 | **Batching Strategy**: Group related changes into batches of 3-10 changes. This makes debugging manageable while maintaining efficiency. Test each batch before moving to the next. |
| 74 | |
| 75 | **Principle: Minimal, Precise Edits** |
| 76 | When implementing tracked changes, only mark text that actually changes. Repeating unchanged text makes edits harder to review and appears unprofessional. Break replacements into: [unchanged text] + [deletion] + [insertion] + [unchanged text]. Preserve the original run's RSID for unchanged text by extracting the `<w:r>` element from the original and reusing it. |
| 77 | |
| 78 | Example - Changing "30 days" to "60 days" in a sentence: |
| 79 | ```python |
| 80 | # BAD - Replaces entire sentence |
| 81 | '<w:del><w:r><w:delText>The term is 30 days.</w:delText></w:r></w:de |