name: office-docs description: Comprehensive document processing for Microsoft Word (.docx) and WPS Office files. Use when Codex needs to work with professional documents for: (1) Creating new documents, (2) Modifying or editing content, (3) Converting between formats, (4) Extracting text and metadata, (5) Troubleshooting document issues, (6) Batch processing documents, or any other Office document tasks.
This skill provides comprehensive tools and workflows for working with Microsoft Word (.docx) and WPS Office documents. It covers creation, editing, conversion, analysis, and troubleshooting of professional documents.
Read document content:
# Use python-docx for .docx files
from docx import Document
doc = Document('document.docx')
text = '\n'.join([paragraph.text for paragraph in doc.paragraphs])
Create new document:
from docx import Document
from docx.shared import Inches
doc = Document()
doc.add_heading('Document Title', 0)
doc.add_paragraph('This is a new paragraph.')
doc.save('new_document.docx')
For .docx files:
- python-docx - Primary library for reading/writing .docx
- docx2txt - Simple text extraction
- docxcompose - Advanced document composition
- docx-mailmerge - Mail merge functionality
For WPS files:
- pywps - WPS file manipulation (when available)
- Conversion to .docx first recommended
For format conversion:
- pandoc - Universal document converter
- libreoffice - Office suite for conversion
- unoconv - Universal office converter
Document conversion:
# Convert .docx to PDF
libreoffice --headless --convert-to pdf document.docx
# Convert .docx to text
pandoc document.docx -o document.txt
# Batch convert WPS to .docx
for file in *.wps; do libreoffice --headless --convert-to docx "$file"; done
Document analysis:
# Extract metadata
exiftool document.docx
# Check file integrity
file document.docx
When creating new documents:
See CREATION.md for detailed patterns.
When modifying existing documents:
See EDITING.md for detailed patterns.
When converting between formats:
See CONVERSION.md for detailed patterns.
Symptoms: Won't open, error messages, missing content
Solutions:
- Try opening in different application
- Use recovery mode in Word/WPS
- Extract content with python-docx ignoring errors
- Convert to different format and back
See TROUBLESHOOTING.md for detailed recovery procedures.
Symptoms: Wrong fonts, broken layout, missing styles
Solutions: - Check style definitions - Verify font availability - Use template-based approach - Simplify complex formatting
Symptoms: Different appearance in Word vs WPS, missing features
Solutions: - Stick to common features - Test in both applications - Use standard formats - Provide alternative versions
Batch processing:
import os
from docx import Document
def process_documents(folder_path):
for filename in os.listdir(folder_path):
if filename.endswith('.docx'):
doc_path = os.path.join(folder_path, filename)
process_single_document(doc_path)
Template-based generation:
from docx import Document
def generate_from_template(template_path, data):
doc = Document(template_path)
# Replace placeholders with data
for paragraph in doc.paragraphs:
for key, value in data.items():
if f'{{{{ {key} }}}}' in paragraph.text:
paragraph.text = paragraph.text.replace(f'{{{{ {key} }}}}', value)
return doc
Extract statistics:
def analyze_document(doc_path):
doc = Document(doc_path)
stats = {
'paragraphs': len(doc.paragraphs),
'tables': len(doc.tables),
'images': len(doc.inline_shapes),
'sections': len(doc.sections),
'styles': len(doc.styles)
}
return stats
Check formatting consistency:
def check_formatting(doc):
issues = []
for i, para in enumerate(doc.paragraphs):
if para.style.name == 'Normal' and para.text.strip():
# Check for inconsistent formatting
if len(para.runs) > 1:
issues.append(f"Paragraph {i}: Multiple runs in Normal style")
return issues
import shutil
import os
def backup_document(filepath):
backup_path = filepath + '.backup'
shutil.copy2(filepath, backup_path)
return backup_path
try:
doc = Document(filepath)
except Exception as e:
print(f"Error opening {filepath}: {e}")
# Try alternative methods
return extract_text_fallback(filepath)
For detailed information on specific topics, consult these reference files:
Available scripts in the scripts/ directory:
extract_text.py - Extract text from .docx filesconvert_format.py - Convert between document formatsbatch_process.py - Process multiple documentsdocument_stats.py - Generate document statisticsrepair_document.py - Attempt to repair corrupted documentsRun scripts with appropriate parameters:
python scripts/extract_text.py input.docx output.txt
If you encounter issues not covered in this skill:
Remember: When in doubt, create a backup and work on a copy.
这个 Skill 功能较为全面,能处理 Word 和 WPS 文档的多种操作。文档结构清晰,提供了实用的代码示例和工作流程指导。但内容不完整,引用了多个未实现的文档,部分内容被截断。对于需要高级功能的用户来说可能不够用,建议等待版本更新后再使用。