dof
dof scans a directory recursively for common document files and maintains an Excel index (a "treasure map").
Installation
pip install treasure-map
Installed as treasure-map, used as dof. The distribution on PyPI is named
treasure-map; the import package and the console script are both still dof.
pip install dof will not work — that name is unavailable on PyPI.
| Name | |
|---|---|
| Install | pip install treasure-map |
| Command | dof (or python -m dof) |
| Import | import dof |
To install from a source checkout:
pip install -e ".[dev]"
CLI
# Scan current directory and write ./treasure_map.xlsx
dof
# Scan a specific directory
dof -d /path/to/root
# Choose output filename
dof -d . -o my_map.xlsx
# Use a SharePoint/OneDrive base URL for hyperlinks
export DOF_SHAREPOINT_BASE_URL="https://example.sharepoint.com/sites/Team/Shared%20Documents"
dof -d .
# Preview changes without writing (dry run)
dof --dry-run
# Output as JSON or CSV instead of Excel
dof --format json
dof --format csv
# Keep rows for deleted files (default: remove them)
dof --keep-missing
# Don't exit non-zero when broken links are found
dof --keep-missing --no-fail-on-broken
# Turn off move tracking (a move becomes a deletion plus a new document)
dof --no-detect-moves
# Disable progress indicator
dof --no-progress
# Verbose logging
dof -v # info level
dof -vv # debug level
CLI Options
| Option | Description |
|---|---|
-d, --dir PATH |
Directory to scan (default: current directory) |
-o, --output PATH |
Output filename (default: treasure_map.xlsx) |
--format [xlsx|json|csv] |
Output format (default: xlsx) |
--dry-run |
Show what would change without writing files |
--keep-missing |
Keep rows for files that no longer exist (default: remove) |
--no-detect-moves |
Treat a moved or renamed file as a deletion plus a new document (default: track the move) |
--no-fail-on-broken |
Exit 0 even when broken links are found (default: exit 2) |
--sharepoint-base URL |
Base SharePoint/OneDrive URL for hyperlinks |
--progress / --no-progress |
Show/hide progress during scan (default: show) |
-v, --verbose |
Enable info-level logging |
-vv, --very-verbose |
Enable debug-level logging |
-h, --help |
Show help message |
--version |
Show version |
Output Columns
| Column | Description |
|---|---|
| File Name | Name of the document file |
| File Type | Document type (PDF, Word, Excel, etc.) |
| Description | User-editable notes (preserved across updates) |
| Date Found | First time the file was discovered (immutable) |
| Last Seen | Most recent scan where the file was present |
| Link | Clickable hyperlink to the file |
| Version | Starts at 1.0; increments when content changes |
| Location | Path relative to the scan root (POSIX-style) |
| Status | OK, Moved or Broken for the current scan |
| Previous Location | Where the file was before it last moved (blank if it never has) |
Tracking Moved Files
Rows used to be keyed solely on Location, so reorganising a folder tree made every
moved document look like a deletion plus a brand-new file. Its Date Found, your
hand-written Description and its whole Version history were all lost, and the
hyperlinks in any copy of the map you had already circulated pointed at nothing.
dof now recognises a moved or renamed document and relinks the existing row to its new
path, keeping Date Found, Description and Version intact. Matching happens in
three tiers:
- Identical content. A discovered file whose SHA-256 matches an orphaned row is
the same document: the row is relinked and its
Versionis left alone. This covers plain moves and renames. - Same name, size and file type. Treated as the same document, moved and edited:
the row is relinked and its
Versionis bumped (e.g.1.0→1.1). - No match. The file is treated as genuinely new, and the old row is pruned (or
marked
Brokenunder--keep-missing).
Empty files never pair. Every zero-byte file shares the same hash and the same size, so neither signal is evidence of anything; such files always fall through to tier 3. A file dof could not hash or measure is likewise never paired — it degrades to "deleted plus new" rather than guessing.
Tier 2 is a probabilistic match. Two unrelated files of the same name, size and type —
say an invoice.docx template in two departments, one removed and one added in the same
scan — can be paired wrongly. Previous Location makes that visible and auditable.
If that ever happens to you, --no-detect-moves is the escape hatch. It turns matching
off altogether and restores the pre-feature behaviour: a moved file is reported as a
deletion plus a new document, losing Date Found, Description and Version but never
pairing two unrelated files. Move tracking is on by default.
Worked example
Say you start with this tree and run dof:
docs/
proposal.pdf
notes/meeting.docx
| File Name | Version | Location | Status | Previous Location |
|---|---|---|---|---|
| proposal.pdf | 1.0 | docs/proposal.pdf | OK | |
| meeting.docx | 1.0 | docs/notes/meeting.docx | OK |
You add a description to proposal.pdf in Excel, then restructure: docs/ becomes
archive/2025/, and you edit meeting.docx while you are there (without changing its
byte length). Rescanning gives:
| File Name | Version | Location | Status | Previous Location |
|---|---|---|---|---|
| proposal.pdf | 1.0 | archive/2025/proposal.pdf | Moved | docs/proposal.pdf |
| meeting.docx | 1.1 | archive/2025/notes/meeting.docx | Moved | docs/notes/meeting.docx |
Both descriptions and both Date Found dates survive. proposal.pdf kept version 1.0
(tier 1, content unchanged); meeting.docx went to 1.1 (tier 2, moved and edited).
The CLI reports:
Moved files:
> docs/proposal.pdf -> archive/2025/proposal.pdf
> docs/notes/meeting.docx -> archive/2025/notes/meeting.docx
Previous Location is sticky: once written it stays, so someone holding an older copy of
the map can still work out where a document went. Status, by contrast, describes the
current scan only — it reverts to OK on the next scan that finds the file where the map
says it is.
Broken and Repaired Links
After each scan dof checks that every row's link still resolves, and repairs what it
can before reporting anything as broken. If a row's file was found by the scan but its
stored target no longer resolves, the target is regenerated from the current scan root
(or from the configured SharePoint base). Repaired rows are listed under Repaired
links: and do not affect the exit code — repair is a success, not a failure.
Repaired links:
* reports/q4_summary.pdf
* notes/meeting_2025.docx
The common real-world trigger is a moved root: you rename or reorganise the top-level
folder, or the whole tree turns up at a different absolute path — on another machine, or
after OneDrive re-roots your local copy. Every relative Location is still correct, but
every stored absolute link still points at the old root. Before v0.1.2 dof marked all of
those rows Broken and exited 2. Now a plain re-run repairs them and exits 0:
dof -d /path/to/renamed/tree
Repair is targeted: a link that already resolves is never regenerated, so a hyperlink you edited by hand is not clobbered.
Broken therefore means dof cannot repair this — not merely that the link is stale.
A broken row is marked Status = Broken, highlighted red in the workbook, and listed
under Broken links: in the CLI output. dof then exits with code 2.
This mostly arises with --keep-missing, which deliberately keeps rows for files that
are gone — those files really are missing, so there is nothing to repair. Pass
--no-fail-on-broken to keep the marking and the report but exit 0:
dof --keep-missing --no-fail-on-broken
Link checking never touches the network. A file:// target resolves when the path exists
on disk; a SharePoint target resolves when its path, taken relative to the configured base
URL, matches something found in this scan. A hand-pasted http(s):// link with no
--sharepoint-base configured is always treated as resolvable, because dof has no offline
way to judge it — so a genuinely dead SharePoint URL for a file that is present locally
still reports OK.
Exit Codes
| Code | Meaning |
|---|---|
| 0 | Success |
| 1 | Error (see the message) |
| 2 | Unrepairable broken links present (suppress with --no-fail-on-broken) |
Repaired links never cause a non-zero exit.
Ignore Patterns
Create a .treasureignore file in the scan root to exclude files using gitignore-style patterns:
# Ignore all .tmp files
*.tmp
# Ignore build directory
build/
# Ignore specific file
secret.pdf
# But keep this one
!important.pdf
Supported File Types
Office: .doc, .docx, .xls, .xlsx, .ppt, .pptx, .rtf
Text: .txt, .md, .rst, .csv, .json, .yaml, .xml, .toml
PDF: .pdf
Other: .odt, .ods, .odp, .pages, .numbers, .key, .epub, .tex
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file treasure_map-0.1.3.tar.gz.
File metadata
- Download URL: treasure_map-0.1.3.tar.gz
- Upload date:
- Size: 82.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d0a3ca43a8517d80bd018c0fb922a97e11b6496ec58dc3a97394b76979532c22
|
|
| MD5 |
4feb225e564ad6677d3c649ef8071cf1
|
|
| BLAKE2b-256 |
90892c659ffbc858f33ab86caa722040a7e794d9e1bd3b31a83c52f39dd6ce1c
|
File details
Details for the file treasure_map-0.1.3-py3-none-any.whl.
File metadata
- Download URL: treasure_map-0.1.3-py3-none-any.whl
- Upload date:
- Size: 23.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.12.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d18a2bc0bf988f9fe923221bdfbd39275bdfd909b71d6f1796053b2b82fc0d52
|
|
| MD5 |
013a70f90f546984e8040a2f2ccfae81
|
|
| BLAKE2b-256 |
22fe6d89e08ba341c9f492bb8c069748e186edb4f913a5ed35eaffdfd5d0f4cd
|