A C++17 shared library that extracts every piece of comment metadata from .docx files — text, authors, dates, reply threads, anchor text, and resolution status — with full Python bindings via pybind11.
Since v1.2 it also turns those comments into spreadsheets, DataFrames and JSON, and ships a command-line tool so you can use it without writing any Python.
- What it does
- What's new in v1.2
- Quick start — Python
- Quick start — command line
- Quick start — C++
- Installation
- Exporting comments
- Command-line guide
- Python API reference
- C++ API reference
- Architecture
- Performance
- Testing
- Changelog
- License
A .docx file is a ZIP archive containing XML parts defined by the OOXML standard. Comments are spread across up to four of those parts, each requiring a different parsing strategy:
| Part | Content | Parse method |
|---|---|---|
word/comments.xml |
Core comment data (id, author, date, text) | DOM — always small |
word/commentsExtended.xml |
Reply threading, done flag (OOXML 2016+) |
SAX streaming |
word/commentsIds.xml |
Para-ID cross-reference (fallback) | SAX streaming |
word/document.xml |
Anchor text via commentRangeStart/End |
SAX streaming — can be very large |
docx_comment_parser opens the ZIP without decompressing it fully, inflates each part on demand, parses it, and discards the raw bytes. The result is a fully resolved CommentMetadata object for every comment in the document, with reply chains linked by id and anchor text extracted from the document body.
What you get per comment:
- Identity:
id,author,initials,date(ISO-8601 string) - Content:
text(full plain-text body, XML entities decoded),paragraph_style - Anchoring:
referenced_text— the exact document text the comment is attached to - Threading:
is_reply,parent_id,replieslist,thread_idschain - Resolution:
doneflag fromcommentsExtended.xml
Everything from v1.1 still works exactly as before. v1.2 adds two things on top.
1. You can get your comments as a table.
Before, you had to loop over comment objects and build your own rows. Now one method call gives you a spreadsheet, a DataFrame, or JSON:
parser.to_dataframe() # pandas
parser.to_polars() # polars
parser.export_csv("out.csv") # spreadsheet — no extra packages needed
parser.export_json("out.json") # JSON — no extra packages needed2. You can use it from a terminal, without writing Python.
docx-comments parse report.docx # see the comments
docx-comments stats report.docx # who commented, how much is done
docx-comments unresolved report.docx # what's still open
docx-comments export report.docx --csv -o comments.csv
docx-comments batch ./documents # a whole folder at onceNothing got heavier. Installing the package still pulls in zero dependencies. pandas, polars and the CLI tools are optional extras you opt into. The parser itself is unchanged and just as fast — see Performance.
Two long-standing bugs were fixed along the way; both are described in the Changelog.
The compiled C++ module moved from being the whole package to sitting inside it, at docx_comment_parser._core. This is invisible in normal use — import docx_comment_parser as dcp and dcp.DocxParser() behave identically. The only code affected is anything that imported the private extension file by path, which was never a supported thing to do.
import docx_comment_parser as dcp
parser = dcp.DocxParser()
parser.parse("report.docx")
# Print every comment
for c in parser.comments():
prefix = " ↳ [reply]" if c.is_reply else f"[{c.id}]"
print(f"{prefix} {c.author} ({c.date[:10]}): {c.text[:80]}")
if c.referenced_text:
print(f" anchored to: \"{c.referenced_text[:60]}\"")[0] Alice (2026-01-15): This sentence needs rephrasing for clarity and conciseness.
anchored to: "The methodology employed in this study is fundamentally flaw"
↳ [reply] Bob (2026-01-16): Agreed. Suggest: "This sentence requires revision."
[2] Alice (2026-01-17): Please verify the statistical analysis in section 3 & 4.
anchored to: "Results in section 3 and 4 show p < 0.05."
The same parser can hand you the whole document as rows:
import docx_comment_parser as dcp
parser = dcp.DocxParser()
parser.parse("report.docx")
# A spreadsheet you can open in Excel — needs nothing extra installed.
parser.export_csv("comments.csv")
# A pandas DataFrame — needs `pip install docx-comment-parser[pandas]`.
df = parser.to_dataframe()
print(df[["author", "text", "resolved"]].head()) author text resolved
0 Alice This sentence needs rephrasing for clari… False
1 Bob Agreed. Suggest: "This sentence requires… True
2 Alice Please verify the statistical analysis i… False
Because it is a real DataFrame, ordinary pandas works on it:
# Who has the most open comments?
open_by_author = df[~df["resolved"]].groupby("author").size()
# How many comments mention security?
security = df[df["text"].str.contains("security", case=False)]Install the CLI extra once:
pip install "docx-comment-parser[cli]"Then look at a document without writing any code:
docx-comments parse report.docx Comments — report.docx
ID Author Date St Comment Anchored to
────────────────────────────────────────────────────────────────────────────────────
0 Alice 2026-01-15 09:12 ○ This sentence needs rephrasing… The methodology…
1 Bob 2026-01-16 11:03 ✓ ↳ Agreed. Suggest: "This sen…
2 Alice 2026-01-17 14:40 ○ Please verify the statistical… Results in sec…
3 comment(s) 1 resolved 2 open
○ means open, ✓ means resolved, and ↳ marks a reply.
On a terminal that cannot display those characters — a stock Windows console, for instance — the same table prints with plain ASCII (open / done / >) instead. Nothing is lost and nothing crashes; the tool checks what your terminal can handle and adapts.
A few more things you can do:
# Only Alice's comments
docx-comments parse report.docx --author alice
# Only comments that mention "security", anywhere in the comment or the text it points at
docx-comments parse report.docx --contains security
# Turn a folder of documents into one spreadsheet
docx-comments batch ./reviews -o all_comments.csvThe full command reference is in the Command-line guide.
#include "docx_comment_parser.h"
#include <iostream>
int main() {
docx::DocxParser parser;
parser.parse("report.docx");
for (const auto& c : parser.comments()) {
std::cout << "[" << c.id << "] "
<< c.author << ": "
<< c.text.substr(0, 80) << "\n";
if (!c.referenced_text.empty())
std::cout << " anchored to: \"" << c.referenced_text << "\"\n";
}
const auto& s = parser.stats();
std::cout << "\n" << s.total_comments << " comment(s), "
<< s.unique_authors.size() << " author(s)\n";
}The base package has no dependencies at all. Optional features live behind extras, so you only install what you use:
pip install docx-comment-parser # parser + CSV/JSON export. Zero dependencies.
pip install "docx-comment-parser[pandas]" # + to_dataframe()
pip install "docx-comment-parser[polars]" # + to_polars()
pip install "docx-comment-parser[cli]" # + the docx-comments command
pip install "docx-comment-parser[all]" # everything above| Extra | Adds | Gives you |
|---|---|---|
| (none) | — | DocxParser, BatchParser, export_csv(), export_json(), to_dict(), to_json() |
pandas |
pandas ≥ 2.0 | to_dataframe() |
polars |
polars ≥ 1.0 | to_polars() |
cli |
typer, rich | the docx-comments terminal command |
all |
all of the above | everything |
If you call a method whose extra is missing, you get a message telling you exactly what to install rather than an obscure ImportError:
ImportError: pandas is required for this export but is not installed.
Install it with: pip install docx-comment-parser[pandas]
# 1. Install system dependencies
sudo apt install build-essential g++ cmake zlib1g-dev # Debian/Ubuntu
brew install cmake zlib # macOS
# 2. Install the Python build dependency
pip install pybind11
# 3a. Build the Python extension in-place (for development)
python setup.py build_ext --inplace
# 3b. OR install permanently into the current environment
pip install .Verify:
python -c "import docx_comment_parser; print('OK')"docx_comment_parser bundles a self-contained DEFLATE inflate implementation (vendor/zlib/zlib.h). No external zlib install is needed on MSVC — pybind11 is the only dependency.
# 1. Open "Developer Command Prompt for VS 2022" (or run vcvarsall.bat x64)
# 2. Install the only required Python dependency
pip install pybind11
# 3. Build
python setup.py build_ext --inplaceVerify:
python -c "import docx_comment_parser; print('OK')"The compiler invocation will include -Ivendor and no /link zlib.lib:
cl.exe /c /nologo /O2 /std:c++17 /DDOCX_BUILDING_DLL
-Iinclude -Ivendor -I<pybind11\include> ...
/Tpsrc/zip_reader.cpp ...
link.exe ... /OUT:docx_comment_parser.cp314-win_amd64.pyd
# Inside an MSYS2 MINGW64 shell
pacman -S mingw-w64-x86_64-gcc mingw-w64-x86_64-cmake \
mingw-w64-x86_64-zlib mingw-w64-x86_64-python \
mingw-w64-x86_64-python-pip
pip install pybind11
python setup.py build_ext --inplaceIf you need the C++ .so/.dll without Python bindings:
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)CMake build options:
| Option | Default | Effect |
|---|---|---|
BUILD_PYTHON_BINDINGS |
ON |
Compile the pybind11 extension |
BUILD_TESTS |
ON |
Build and register the test suite with CTest |
CMAKE_BUILD_TYPE |
Release |
Debug / Release / RelWithDebInfo |
parser.comments() gives you comment objects shaped like the OOXML file format. That is the right shape for reading one comment at a time, but the wrong shape for a spreadsheet: reply links use -1 to mean "no parent", "resolved" is called done, dates are raw text, and nothing records which file a comment came from.
The export layer flattens all of that into plain rows. One comment = one row. Same columns every time.
Every method works on any parsed document:
parser = dcp.DocxParser()
parser.parse("report.docx")
rows = parser.to_comments() # list of Comment objects
df = parser.to_dataframe() # pandas DataFrame [pandas]
pf = parser.to_polars() # polars DataFrame [polars]
dicts = parser.to_dict() # list of plain dicts
text = parser.to_json() # JSON string
parser.export_csv("comments.csv") # write a CSV file
parser.export_json("comments.json") # write a JSON fileexport_csv and export_json return the path they wrote, and create missing folders for you:
path = parser.export_csv("reports/2026/q1/comments.csv") # folders created
print(f"Wrote {path}")| Column | Type | What it is |
|---|---|---|
comment_id |
int | The comment's id in the document |
parent_id |
int or empty | The comment this one replies to. Empty for a top-level comment |
author |
str | Who wrote it |
initials |
str | Their initials, as Word recorded them |
date |
str | The timestamp exactly as stored in the file |
date_parsed |
datetime | The same timestamp as a real date you can sort and filter on |
text |
str | The comment itself |
referenced_text |
str | The document text the comment points at |
paragraph_style |
str | Word style of the comment's first paragraph |
resolved |
bool | Whether it has been marked resolved |
is_reply |
bool | Whether it is a reply to another comment |
thread_depth |
int | 0 for a top-level comment, 1 for a reply, 2 for a reply to a reply… |
document_name |
str | Which file it came from |
root_id |
int | The id of the first comment in this conversation |
reply_count |
int | How many direct replies it has |
para_id, para_id_parent |
str | Word's internal paragraph ids |
range_start_para_id, range_end_para_id |
str | Ids marking where the comment is anchored |
paragraph_index |
int | Which paragraph in the document it is attached to (-1 if unknown) |
run_index |
int | Which run inside that paragraph (-1 if unknown) |
The first thirteen are what most people use. The rest carry the low-level anchoring detail through, so exporting never loses information compared with reading parser.comments() directly.
date is the untouched string from the file. date_parsed is that string turned into a real datetime. You get both because they fail differently: if Word wrote something unusual, date_parsed becomes empty but date still shows you exactly what was in the document. No data is ever silently lost, and a single odd timestamp cannot break a 10,000-comment export.
df["date_parsed"].dt.month # works like any datetime column
df[df["date_parsed"] > "2026-01-01"] # filter by datefilter_comments applies the same rules the CLI uses. Every argument is optional and they combine with AND:
from docx_comment_parser import DocxParser
from docx_comment_parser.filters import filter_comments
from docx_comment_parser.exporters import export_csv
parser = DocxParser()
parser.parse("report.docx")
open_security_notes = filter_comments(
parser.to_comments(),
contains="security", # in the comment OR the text it points at
resolved=False, # only unresolved
)
export_csv(open_security_notes, "security_todo.csv")| Argument | Effect |
|---|---|
author="alice" |
Author contains "alice", ignoring case. Matches "Alice Smith" |
contains="security" |
The word appears in the comment text or in the text it points at |
resolved=True / False / None |
Only resolved / only open / both |
threads_only=True |
Only comments that are part of a conversation, dropping standalone notes |
BatchParser parses files in parallel and exports them as one combined table. The document_name column tells you which file each row came from:
import glob
import docx_comment_parser as dcp
batch = dcp.BatchParser(max_threads=0) # 0 = use every CPU core
batch.parse_all(glob.glob("reviews/*.docx"))
df = batch.to_dataframe()
print(df.groupby("document_name").size()) # comments per file
batch.export_csv("all_reviews.csv")Files that fail to parse do not stop the run. They are reported separately and skipped by the export:
for path, message in batch.errors().items():
print(f"Could not read {path}: {message}")
print(batch.parsed_files()) # only the files that workedThe methods above are thin wrappers. If you have built your own list of comments, the underlying functions take it directly:
from docx_comment_parser.exporters import (
to_dataframe, to_polars, to_dict, to_json, export_csv, export_json,
)
mine = [c for c in parser.to_comments() if c.author == "Alice"]
to_dataframe(mine)
export_csv(mine, "alice.csv")export_csv writes UTF-8. If you plan to open the file by double-clicking it in Excel on Windows, ask for the byte-order mark so accented names survive:
parser.export_csv("comments.csv", encoding="utf-8-sig")
parser.export_csv("comments.csv", delimiter=";") # for locales where Excel expects ;to_json always produces valid JSON with dates as ISO-8601 strings, so it can be posted to an API or read back with json.loads without a custom decoder.
Install with pip install "docx-comment-parser[cli]", then run docx-comments --help. Every command has its own --help too.
The command exists even without the extra installed — it just tells you how to install it instead of crashing.
docx-comments parse report.docx
docx-comments parse report.docx --author alice --unresolved
docx-comments parse report.docx --limit 20docx-comments stats report.docx╭─ report.docx ─────────────────╮
│ Total comments 42 │
│ Root comments 18 │
│ Replies 24 │
│ Resolved 31 │
│ Unresolved 11 │
│ Unique authors 4 │
│ Earliest comment 2026-01-15 │
│ Latest comment 2026-02-02 │
╰───────────────────────────────╯
By author
Author Comments Resolved Open Resolution rate
──────────────────────────────────────────────────────
Alice 19 15 4 79%
Bob 12 9 3 75%
Carol 11 7 4 64%
Prints the open comments and exits with status 1 if there are any. That makes it usable as a gate in a script or CI job:
docx-comments unresolved spec.docx || echo "Review is not finished yet"Exit code 0 means nothing is left open.
docx-comments export report.docx --csv -o comments.csv
docx-comments export report.docx --json -o comments.jsonWith no -o, the data goes to standard output so it can be piped:
docx-comments export report.docx --json | jq '.[] | select(.resolved == false) | .author'If you give -o a filename, the format is inferred from the extension, so --csv / --json are optional:
docx-comments export report.docx -o comments.csv # CSV, inferreddocx-comments batch ./reviews
docx-comments batch ./reviews --recursive --threads 8
docx-comments batch ./reviews -o all_comments.csvPrints one row per file, then a total. Word's ~$name.docx lock files are ignored. Unreadable files are listed at the end and the command exits 1, but every readable file is still processed and exported.
--author, --contains, --resolved, --unresolved and --threads-only work the same way on parse, export and batch:
| Flag | Meaning |
|---|---|
--author NAME, -a |
Author contains NAME, ignoring case |
--contains TEXT, -c |
TEXT appears in the comment or the text it points at |
--resolved |
Only resolved comments |
--unresolved |
Only open comments |
--threads-only |
Only comments that are part of a conversation |
--limit N, -n |
Show at most N comments (parse, unresolved) |
--resolved and --unresolved together is an error, since nothing could match.
| Code | Meaning |
|---|---|
0 |
Success |
1 |
The file could not be read, or unresolved found open comments, or batch hit an unreadable file |
2 |
The command line itself was wrong |
import docx_comment_parser as dcpSingle-file parser. Non-copyable, movable. Can be reused across multiple calls to parse().
Parses a .docx file and populates all results. Replaces any previous results from an earlier call.
parser = dcp.DocxParser()
parser.parse("report.docx")Raises DocxFileError if the file cannot be opened or is not a valid ZIP archive.
Raises DocxFormatError if the OOXML structure is malformed.
Files without any comments parse successfully and return an empty list from comments().
Returns all comments sorted ascending by id.
for c in parser.comments():
print(f"#{c.id:3d} {c.author:20s} {c.text[:60]}")Looks up a single comment by its w:id. Returns None if not found.
c = parser.find_by_id(3)
if c is not None:
print(c.author, "—", c.text)Returns all comments whose author field exactly matches the given string (case-sensitive). The author string is taken directly from the w:author XML attribute.
for c in parser.by_author("Alice"):
status = "✓" if c.done else "○"
print(f" {status} [{c.date[:10]}] {c.text[:70]}")Returns only the top-level (non-reply) comments in document order.
for root in parser.root_comments():
n = len(root.replies)
print(f"Thread #{root.id}: {n} repl{'y' if n == 1 else 'ies'}")Returns the full reply chain for a given root comment, starting with the root itself, in chronological order.
for c in parser.thread(0):
indent = " " if c.is_reply else ""
print(f"{indent}[{c.id}] {c.author}: {c.text}")[0] Alice: This sentence needs rephrasing for clarity and conciseness.
[1] Bob: Agreed. Suggest: "This sentence requires revision."
Returns aggregate statistics computed during the last parse() call.
s = parser.stats()
print(f"File : {s.file_path}")
print(f"Comments : {s.total_comments} total "
f"({s.total_root_comments} root, {s.total_replies} replies)")
print(f"Resolved : {s.total_resolved}")
print(f"Authors : {', '.join(s.unique_authors)}")
print(f"Date range: {s.earliest_date[:10]} → {s.latest_date[:10]}")File : report.docx
Comments : 3 total (2 root, 1 replies)
Resolved : 1
Authors : Alice, Bob
Date range: 2026-01-15 → 2026-01-17
Added in v1.2. All of them operate on the currently parsed document. See Exporting comments for the full column list and examples.
| Method | Returns | Needs |
|---|---|---|
to_comments() |
list[Comment] |
— |
to_dict() |
list[dict] |
— |
to_json(indent=2) |
str |
— |
export_json(path, indent=2) |
Path written |
— |
export_csv(path, encoding="utf-8", delimiter=",") |
Path written |
— |
to_dataframe() |
pandas.DataFrame |
[pandas] extra |
to_polars() |
polars.DataFrame |
[polars] extra |
parser.parse("report.docx")
parser.to_dataframe() # a table
parser.export_csv("comments.csv") # a spreadsheetProcesses many files in parallel using a thread pool. The Python GIL is released during parse_all, so CPU-bound threads are not blocked.
bp = dcp.BatchParser(max_threads=0) # 0 = one thread per CPU coreParses all files. Files that raise errors are captured in errors() rather than propagating as exceptions, so one bad file does not abort the batch.
Returns the parsed comments for a specific file.
Returns statistics for a specific file.
Returns {file_path: error_message} for every file that failed.
for path, msg in bp.errors().items():
print(f"FAILED {path}: {msg}")Frees the in-memory results for one file. Call this as soon as you have finished processing a file to keep peak memory low when working with large batches.
Frees results for all files.
Complete batch example:
import docx_comment_parser as dcp
import glob, json
files = glob.glob("/documents/**/*.docx", recursive=True)
bp = dcp.BatchParser(max_threads=0)
bp.parse_all(files)
summary = []
for path in files:
if path in bp.errors():
print(f"SKIP {path}: {bp.errors()[path]}")
continue
s = bp.stats(path)
summary.append({
"file": path,
"comments": s.total_comments,
"authors": s.unique_authors,
"resolved": s.total_resolved,
})
bp.release(path) # free this file's memory immediately
print(json.dumps(summary, indent=2))Added in v1.2. The files that parsed successfully and still hold results, sorted. Files that failed and files you have already release()d are not listed.
bp.parse_all(["a.docx", "b.docx", "broken.docx"])
bp.parsed_files() # ['a.docx', 'b.docx']Added in v1.2. Same methods as DocxParser, but they combine every parsed file into one table, with the document_name column identifying the source. Each takes an optional file_paths argument to restrict the export; the default is every successfully parsed file.
bp.parse_all(glob.glob("reviews/*.docx"))
bp.to_dataframe() # all files, one table
bp.to_dataframe(file_paths=["a.docx"]) # just one
bp.export_csv("all_reviews.csv")Added in v1.2. Comment is the flat, tabular version of CommentMetadata returned by to_comments() and used as the row type by every exporter. The full column table is in Exporting comments.
The differences from CommentMetadata are deliberate, and they are what make it table-friendly:
CommentMetadata |
Comment |
Why |
|---|---|---|
id |
comment_id |
Unambiguous as a column heading |
parent_id == -1 |
parent_id is None |
A missing value, not a magic number |
done |
resolved |
Says what it means |
date (string only) |
date and date_parsed |
Keeps the original, adds a usable datetime |
| — | thread_depth, root_id, reply_count |
Conversation position, computed for you |
| — | document_name |
Which file the row came from |
from docx_comment_parser import Comment, FIELD_NAMES
FIELD_NAMES # the canonical column order, shared by every exporter
comment.to_dict() # one row as a plain dictAll fields are read-only. Available in both Python and C++.
| Field | Type | Description |
|---|---|---|
id |
int |
w:id attribute. Unique within the document. |
author |
str |
w:author — display name as set in Word. |
date |
str |
w:date — ISO-8601 string exactly as stored in XML, e.g. "2026-01-15T09:00:00Z". Not parsed into a date object. |
initials |
str |
w:initials — author abbreviation shown in the comment balloon. |
text |
str |
Full plain-text body of the comment. XML character entities are decoded: & → &, < → <, > → >, " → ", ' → ', numeric references → UTF-8. |
paragraph_style |
str |
Style name of the first paragraph inside the comment (e.g. "CommentText"). Empty string if not set. |
referenced_text |
str |
The document text that the comment is anchored to, extracted from the commentRangeStart / commentRangeEnd region in word/document.xml. Truncated to 240 bytes at a UTF-8 boundary. Empty if the range spans no text runs or the file has no word/document.xml. |
is_reply |
bool |
True if this comment is a threaded reply. Requires word/commentsExtended.xml to be present. |
parent_id |
int |
id of the parent comment. -1 for root (non-reply) comments. |
replies |
list[CommentRef] |
Direct child replies, populated on the parent comment. Empty on reply comments. |
thread_ids |
list[int] |
Ordered list of all ids in the full reply chain. Populated only on root comments. Use parser.thread(root_id) to retrieve the full objects. |
done |
bool |
True if the comment has been marked resolved in Word. Sourced from commentsExtended.xml. False when that file is absent. |
para_id |
str |
OOXML 2016+ paragraph ID (w14:paraId). Used internally for thread resolution. |
para_id_parent |
str |
Parent paragraph ID string before numeric id resolution. |
paragraph_index |
int |
0-based paragraph position in the document body. -1 if not determined. |
run_index |
int |
0-based run position within the paragraph. -1 if not determined. |
| Field | Type | Description |
|---|---|---|
id |
int |
id of the reply comment. |
author |
str |
Author of the reply. |
date |
str |
ISO-8601 date of the reply. |
text_snippet |
str |
First 120 characters of the reply text. |
Both CommentMetadata and DocumentCommentStats expose a to_dict() method that returns all fields as a plain Python dict.
import json
data = [c.to_dict() for c in parser.comments()]
print(json.dumps(data, indent=2, ensure_ascii=False))| Field | Type | Description |
|---|---|---|
file_path |
str |
Path passed to parse(). |
total_comments |
int |
Total comments including replies. |
total_root_comments |
int |
Top-level (non-reply) comments. |
total_replies |
int |
Reply comments. Equal to total_comments - total_root_comments. |
total_resolved |
int |
Comments with done=True. |
unique_authors |
list[str] |
Sorted list of distinct author names. |
earliest_date |
str |
ISO-8601 date string of the oldest comment. |
latest_date |
str |
ISO-8601 date string of the most recent comment. |
| Exception | Inherits from | Raised when |
|---|---|---|
dcp.DocxFileError |
DocxParserError, OSError |
File not found, permission denied, or not a valid ZIP archive. |
dcp.DocxFormatError |
DocxParserError, ValueError |
Valid ZIP but required OOXML parts are missing or structurally invalid. |
dcp.DocxParserError |
RuntimeError |
Base class — catches both of the above with a single handler. |
try:
parser.parse("report.docx")
except dcp.DocxFileError as e:
print(f"Cannot open file: {e}")
except dcp.DocxFormatError as e:
print(f"Not a valid .docx: {e}")Each exception is catchable by its own type, by DocxParserError, and by the matching builtin — so all four of these work:
except dcp.DocxFileError: ... # the specific error
except dcp.DocxParserError: ... # anything this library raises
except OSError: ... # any file problem, from any library
except RuntimeError: ... # the broadest baseFixed in v1.2. Before v1.2 the specific types were unreachable: every failure arrived as
DocxParserError, soexcept dcp.DocxFileErrorsilently never matched. Code that catchesDocxParserError,OSErrororValueErroris unaffected and keeps working.
BatchParser.parse_all() never raises. Failures go into errors() instead:
bp.parse_all(["good.docx", "corrupt.docx", "missing.docx"])
print(bp.errors())
# {'corrupt.docx': 'inflate failed...', 'missing.docx': 'Cannot open file...'}Include the single public header:
#include "docx_comment_parser.h"Link against the shared library:
target_link_libraries(my_app PRIVATE docx_comment_parser)docx::DocxParser parser;
// Parse a file — throws on error
parser.parse("report.docx");
// Iterate all comments (sorted by id)
for (const auto& c : parser.comments()) {
std::cout << "[" << c.id << "] "
<< c.author << ": " << c.text << "\n";
}
// Look up by id — returns nullptr if not found
const docx::CommentMetadata* c = parser.find_by_id(2);
if (c) std::cout << c->text << "\n";
// Filter by author
for (const auto* c : parser.by_author("Alice"))
std::cout << c->text << "\n";
// Top-level comments only
for (const auto* root : parser.root_comments())
std::cout << root->id << " has " << root->replies.size() << " replies\n";
// Full reply thread
for (const auto* c : parser.thread(0)) {
std::string indent = c->is_reply ? " " : "";
std::cout << indent << c->author << ": " << c->text << "\n";
}
// Aggregate statistics
const auto& s = parser.stats();
std::cout << s.total_comments << " comments by "
<< s.unique_authors.size() << " authors\n"
<< "Date range: " << s.earliest_date
<< " – " << s.latest_date << "\n";// 0 = use std::thread::hardware_concurrency()
docx::BatchParser bp(/*max_threads=*/0);
bp.parse_all({"a.docx", "b.docx", "c.docx"});
// Check for failures
for (const auto& [path, msg] : bp.errors())
std::cerr << "Failed: " << path << ": " << msg << "\n";
// Access results per file
for (const auto& c : bp.comments("a.docx"))
std::cout << c.author << ": " << c.text << "\n";
std::cout << bp.stats("a.docx").total_comments << "\n";
// Free memory as you go
bp.release("a.docx");
bp.release_all();try {
parser.parse("report.docx");
} catch (const docx::DocxFileError& e) {
// file not found, not a ZIP
} catch (const docx::DocxFormatError& e) {
// valid ZIP, bad OOXML
} catch (const docx::DocxParserError& e) {
// base class — catches both
}docx_comment_parser/
├── include/
│ ├── docx_comment_parser.h ← public API (the only header consumers include)
│ ├── zip_reader.h ← ZIP/DEFLATE reader interface
│ └── xml_parser.h ← SAX + minimal DOM interface
├── src/
│ ├── docx_parser.cpp ← orchestrates all four OOXML parts → CommentMetadata
│ ├── batch_parser.cpp ← std::thread pool + result map
│ ├── zip_reader.cpp ← memory-mapped ZIP + on-demand inflate
│ └── xml_parser.cpp ← self-contained SAX + DOM, no libxml2
├── vendor/
│ └── zlib/
│ └── zlib.h ← vendored DEFLATE + CRC-32 (used on MSVC only)
├── python/
│ └── python_bindings.cpp ← pybind11 module (GIL released during batch)
├── tests/
│ ├── CMakeLists.txt
│ └── test_docx_parser.cpp ← 38 assertions, builds its own .docx in memory
├── CMakeLists.txt
└── setup.py
.docx file (ZIP)
│
▼
ZipReader — memory-mapped — inflate one entry at a time
│
├──▶ word/comments.xml → dom_parse() → CommentMetadata[]
│ id, author, date, initials, text
│
├──▶ word/commentsExtended → sax_parse() → fill is_reply, done, para_id_parent
│
├──▶ word/commentsIds.xml → sax_parse() → fill missing para_ids (fallback)
│
├──▶ resolve_threads() → link parent_id, replies[], thread_ids[]
│
└──▶ word/document.xml → sax_parse() → fill referenced_text per comment
ZIP extraction: the file is memory-mapped (mmap / MapViewOfFile). Each ZIP entry is inflated into a temporary heap buffer, parsed, and the buffer is freed. No two entries' raw bytes are live at the same time.
XML parsing: comments.xml is parsed into a minimal DOM tree (always small — typically < 100 KB). The three other parts are streamed with SAX callbacks; only the data the callbacks accumulate is held in memory, not the raw XML text.
BatchParser: one DocxParser instance per worker thread. Results are stored in a std::unordered_map protected by a mutex. Calling release(path) immediately after consuming a file's results keeps peak memory proportional to max_threads, not to the total batch size.
| Capability | Implementation |
|---|---|
| ZIP parsing | Custom memory-mapped reader (no libzip, no minizip) |
| DEFLATE inflate | System zlib on Linux / macOS / MinGW; vendor/zlib/zlib.h on MSVC |
| XML parsing | Custom SAX + minimal DOM (no libxml2, no expat) |
| Threading | std::thread + std::mutex — C++17 standard library only |
| Python bindings | pybind11 — header-only, build-time dependency only |
Parsing speed is the point of this library, so v1.2 was measured against v1.1.2 to confirm the restructure cost nothing.
Same machine, same documents, runs interleaved so background load affects both equally. Each figure is the best median of five alternating rounds.
| Comments | v1.1.2 | v1.2.0 | Change |
|---|---|---|---|
| 100 | 1.204 ms | 1.166 ms | −3.2% |
| 1,000 | 11.593 ms | 11.594 ms | ±0.0% |
| 10,000 | 125.652 ms | 121.390 ms | −3.4% |
Roughly 80,000–86,000 comments per second, unchanged. The differences are measurement noise, not real gains.
This is the expected result: the parser's C++ code was not touched apart from resetting a stats struct once per parse() call. The export layer is pure Python that runs only when you ask for it, so a program that never calls to_dataframe() pays nothing for its existence.
Measured on the same documents, best of seven runs:
| Comments | parse() |
to_comments() |
to_dataframe() |
to_polars() |
to_json() |
export_csv() |
|---|---|---|---|---|---|---|
| 100 | 1.3 ms | 1.0 ms | 4.3 ms | 1.8 ms | 1.9 ms | 2.7 ms |
| 1,000 | 10.7 ms | 10.1 ms | 17.7 ms | 13.7 ms | 19.4 ms | 22.6 ms |
| 10,000 | 120.6 ms | 117.4 ms | 161.0 ms | 147.4 ms | 209.6 ms | 228.5 ms |
Every export column includes the to_comments() conversion, so the numbers are end-to-end from a parsed document to the finished output.
A 10,000-comment DataFrame takes 161 ms, comfortably inside the 1-second design budget, and cost grows linearly with the number of comments rather than faster. Memory stays proportional too: CSV writing streams row by row, so exporting a large document does not build the whole file in memory first.
These properties are asserted by the test suite, not just measured once — see the perf tests below.
There are two suites: the original C++ one and a Python one added in v1.2. Together they run 254 checks.
pip install "docx-comment-parser[dev]"
pytest # everything
pytest -m "not perf" # skip the slower performance tests
pytest --cov=docx_comment_parser --cov-report=term-missing188 tests, 97% statement coverage — above the 90% project target.
Like the C++ suite, it invents its own fixtures: tests/python/conftest.py builds genuine .docx packages with zipfile and hands them to the real parser. Nothing is mocked, and no sample documents need to exist on disk.
| File | Covers |
|---|---|
test_core_regression.py |
That the v1.1 API still behaves identically — every class, method, field, to_dict() key and exception |
test_models.py |
Field mapping, date parsing, thread depth, malformed input |
test_exporters.py |
pandas, polars, JSON and CSV output, including dtypes, Unicode and empty documents |
test_filters.py |
Filtering rules |
test_cli.py |
Every command, flag, and exit code, through Typer's test runner |
test_performance.py |
Scale and timing budgets (marked perf) |
The regression file is the important one: it exists specifically to prove that moving the compiled module into a package changed nothing a user can see. If it passes, upgrading is safe.
Type checking is enforced too:
mypy # strict mode, cleanThe test suite creates a synthetic .docx file entirely in memory using a minimal ZIP builder and pre-compressed XML fixtures. No sample files need to be present on disk.
# Build and run via CTest
cmake -B build -DBUILD_TESTS=ON -DCMAKE_BUILD_TYPE=Debug
cmake --build build -j$(nproc)
ctest --test-dir build --output-on-failure
# Or run the binary directly for line-by-line output
./build/tests/test_docx_parserExpected output:
Test fixture: /tmp/test_docx_parser_fixture.docx
=== test_basic_parsing ===
=== test_threading ===
=== test_done_flag ===
=== test_anchor_text ===
=== test_by_author ===
=== test_stats ===
=== test_root_comments ===
=== test_batch_parser ===
=== test_missing_file ===
=== test_encoding_utf8_bom ===
=== test_encoding_utf16le ===
=== test_encoding_utf16be ===
=== test_encoding_utf32le ===
=== test_encoding_windows1252 ===
=== test_encoding_iso8859_1 ===
=== test_encoding_numeric_entities ===
──────────────────────────────
Results: 66 passed, 0 failed
The test binary exits with code 0 on full pass, 1 on any failure.
Public API: backward compatible. Existing code needs no changes. The test_core_regression.py suite exists to prove it.
to_dataframe()(pandas),to_polars()(polars),to_dict(),to_json(),export_csv(),export_json()andto_comments()on bothDocxParserandBatchParser.- A new
Commentdataclass: the flat, one-row-per-comment view. Uses__slots__, so 10,000 comments stay cheap. - Computed columns the parser did not previously expose:
thread_depth,root_id,reply_count,document_name, anddate_parsed(a real datetime alongside the untouched original string). filter_comments()for author / keyword / resolved / thread filtering, shared with the CLI.- CSV export streams to disk; DataFrame export builds column-first, keeping a 10,000-comment export at ~161 ms.
parse,stats,export,unresolvedandbatch, built with Typer and Rich.- Filters on every relevant command:
--author,--contains,--resolved,--unresolved,--threads-only,--limit. unresolvedexits1when open comments remain, so it works as a CI gate.exportwrites to stdout by default, so it pipes intojq.
Returns the sorted list of files that parsed successfully and still hold results. This is what lets the batch exporters work without being handed the paths again.
py::register_exception was called with the base class last, and pybind11 tries translators in reverse registration order — so DocxParserError caught every derived type first. Every failure surfaced as DocxParserError, and except dcp.DocxFileError silently never matched, despite being documented.
The three types are now created with PyErr_NewException and a tuple of bases, and dispatched by a single translator with most-derived-first clauses. DocxFileError is now both a DocxParserError and an OSError; DocxFormatError is both a DocxParserError and a ValueError. Code catching any of the old types keeps working; catching the specific types now works too.
DocxParser::Impl::parse returned early when a document had no comments.xml, or an empty one, before reaching compute_stats(). Re-using a parser therefore left the previous document's totals and file_path visible:
parser.parse("has_comments.docx")
parser.parse("no_comments.docx")
parser.stats().file_path # v1.1.2: "has_comments.docx" ← wrong
# v1.2.0: "no_comments.docx"Stats are now reset at the start of every parse().
- The compiled extension moved from the top level to
docx_comment_parser._core, inside a new pure-Python package.import docx_comment_parser as dcpis unchanged. - Optional extras:
[pandas],[polars],[cli],[all],[dev]. The base install still has zero dependencies. - Ships
py.typedand a_core.pyistub;mypy --strictpasses.
- 188 Python tests at 97% coverage, alongside the existing 66 C++ checks.
- Parser throughput verified against v1.1.2 with interleaved A/B runs: no regression (see Performance).
Included multiple text enconding support for a wide range of encondings. Updated unit tests for the new text enconding functionality.
extract_xml_encoding_decl() — scans the XML prolog for encoding="..."
detect_encoding() — BOM detection (UTF-8/16/32 LE/BE) takes precedence, falls back to the XML declaration
utf16_to_utf8() / utf32_to_utf8() — built-in converters (no platform dependency) with correct surrogate-pair handling
Windows path: win_mbcs_to_utf8() via MultiByteToWideChar + WideCharToMultiByte; maps 60+ encoding names to Windows codepage numbers (all Windows-125x, ISO-8859-1..16, Asian, Cyrillic, Thai, OEM codepages)
Linux/macOS path: iconv_convert() via iconv(3) with the same name alias table; handles E2BIG/EILSEQ/EINVAL gracefully
transcode_to_utf8() — public entry point, called at the start of sax_parse() so all parsing paths (DOM and SAX) go through it automatically
CMakeLists.txt — Added find_package(Iconv QUIET) for non-Windows targets; links Iconv::Iconv only when it's a separate library (not built into libc).
test_encoding_utf8_bom — UTF-8 BOM is silently stripped
test_encoding_utf16le / test_encoding_utf16be — BOM-detected UTF-16
test_encoding_utf32le — BOM-detected UTF-32
test_encoding_windows1252 — encoding="windows-1252" with ç, é, ä in content
test_encoding_iso8859_1 — encoding="ISO-8859-1" with é, ñ
test_encoding_numeric_entities — 中 (Chinese) and é (é) references
Public API: unchanged. Existing code does not need modification.
Bug 1 — huff_build: out-of-bounds write in the Huffman symbol table.
The original implementation used canonical code-start values as array indices into syms[]. For the RFC 1951 fixed literal tree, next[9] = 400, so all 112 nine-bit symbols (bytes 144–255, present in any real XML document) were written to syms[400]…syms[511] — well past the 288-element array. This caused silent heap corruption on every inflate call that decoded actual XML text. Synthetic test data with only ASCII symbols (code values < 144, all 8-bit) happened to stay in bounds by coincidence.
Fixed by filling syms[] cumulatively: for each bit-length b in ascending order, all symbols with lens[i] == b are appended in symbol-value order. This exactly matches how huff_decode's index variable navigates the table.
Bug 2 — inflateInit2: wiped the caller's I/O fields.
inflateInit2 called memset(strm, 0, sizeof(*strm)). The real zlib API contract — and the usage in zip_reader.cpp — requires the caller to set next_in, avail_in, next_out, and avail_out before calling inflateInit2. The memset zeroed all four, so every inflate() call received null pointers and zero lengths, returning Z_DATA_ERROR (-3) immediately on the first bit read.
Fixed by only zeroing the fields inflateInit2 actually owns: total_in, total_out, msg, and state.
The PI handler (<?...?>) scanned for the first bare >. A PI whose content contained > would terminate parsing prematurely. Fixed to scan for the correct ?> closing sequence.
vendor/zlib/zlib.h is now a self-contained, header-only DEFLATE decompressor + CRC-32 implementing the exact zlib API surface used by the library. When compiled with MSVC (#ifdef _MSC_VER), zip_reader.cpp defines VENDOR_ZLIB_IMPLEMENTATION and includes this header instead of the system <zlib.h>. On all other platforms the system zlib is used as before.
The result: building the Python extension on Windows now requires only pip install pybind11. No vcpkg, no pre-installed zlib, no additional configuration.
MIT — see LICENSE for the full text.
vendor/zlib/zlib.h is released under MIT-0 (no attribution required).