Recursively scans a directory and outputs a flat, LLM-friendly file tree
Project description
filescan
filescan is a lightweight Python tool for scanning filesystem structures and Python ASTs and exporting them as flat, graph-style representations.
Instead of nested trees, filescan produces stable lists of nodes with parent pointers, making the output:
- easy to post-process
- friendly for CSV / DataFrame / SQL pipelines
- efficient for LLM ingestion and summarization
filescan can operate at two levels:
- filesystem structure (directories & files)
- Python semantic structure (modules, classes, functions, methods)
Both use the same flat graph design and export formats.
Features
Filesystem scanning
- Recursive directory traversal
- Flat node list with explicit
parent_id - Deterministic ordering
- Optional
.gitignore-style filtering - CSV and JSON export
Python AST scanning
- Module, class, function, and method detection
- Nested functions and classes supported
- Stable symbol IDs with parent relationships
- Best-effort function signature extraction
- First-line docstring capture
General
- Shared schema + export model
- Same API for filesystem and AST scanners
- Usable as both a library and a CLI
- Designed for automation, data pipelines, and AI workflows
Installation
pip install filescan
Or for development:
pip install -e .
Quick start (CLI)
Filesystem scan (default)
Scan the current directory and write a CSV:
filescan
Scan a specific directory:
filescan ./data
Export as JSON:
filescan ./data --format json
Specify output base path:
filescan ./data -o out/tree
This generates:
out/
├── tree.csv
└── tree.json
Python AST scan
Scan Python source files and extract symbols:
filescan ./src --ast
Export AST symbols as JSON:
filescan ./src --ast --format json
Custom output path:
filescan ./src --ast -o out/symbols
This generates:
out/
├── symbols.csv
└── symbols.json
Ignore rules (.fscanignore)
filescan supports gitignore-style patterns via pathspec.
Default behavior
- If
--ignore-fileis provided → use it - Otherwise, look for:
./.fscanignore (current working directory)
Ignore rules apply to:
- filesystem scanning
- AST scanning (Python files are skipped if ignored)
Example .fscanignore
.git/
.idea/
build/
dist/
__pycache__/
*.pyc
Output formats
Both filesystem and AST scans produce flat graphs with schema metadata.
Filesystem schema
| Field | Description |
|---|---|
id |
Unique integer ID |
parent_id |
Parent node ID (null for root) |
type |
'd' = directory, 'f' = file |
name |
Base name |
size |
File size in bytes (null for directories) |
CSV example
# id: Unique integer ID for this node
# parent_id: ID of parent node, or null for root
# type: Node type: 'd' = directory, 'f' = file
# name: Base name of the file or directory
# size: File size in bytes; null for directories
id,parent_id,type,name,size
0,,d,data,
1,0,f,example.txt,128
Python AST schema
| Field | Description |
| - | |
| id | Unique integer ID for this symbol |
| parent_id | Parent symbol ID (null for module) |
| kind | module | class | function | method |
| name | Symbol name |
| module_path | File path relative to scan root |
| lineno | Starting line number (1-based) |
| signature | Function or method signature (best-effort) |
| doc | First line of docstring, if any |
Nested functions and classes are represented naturally via parent_id.
Library usage
Filesystem scanner
from filescan import Scanner
scanner = Scanner(
root="data",
ignore_file=".fscanignore",
)
scanner.scan()
scanner.to_csv() # -> ./data.csv
scanner.to_json() # -> ./data.json
Python AST scanner
from filescan import AstScanner
scanner = AstScanner(
root="src",
ignore_file=".fscanignore",
output="out/symbols",
)
scanner.scan()
scanner.to_csv()
scanner.to_json()
Programmatic access
nodes = scanner.scan()
print(len(nodes))
data = scanner.to_dict()
Why filescan?
Most filesystem and code structures are represented as deeply nested trees. While human-readable, they are verbose, hard to query, and inefficient for large-scale processing.
filescan represents both filesystems and codebases as flat graphs because this format is:
-
Compact and token-efficient Flat lists with numeric IDs consume far fewer tokens than recursive trees, making them ideal for LLM context windows.
-
Explicit and unambiguous All relationships are encoded directly via
parent_id. -
Easy to process Flat data works naturally with filtering, joins, grouping, and graph analysis.
This makes filescan especially suitable for:
- SQL / Pandas / DuckDB pipelines
- Static analysis and refactoring tools
- Graph-based code understanding
- LLM-based reasoning and summarization of projects
In short, filescan favors machine-friendly structure over visual trees, enabling scalable, AI-native workflows.
Development
The project uses a src/ layout.
Examples can be run without installation:
python examples/scan_data.py
Or as a module:
python -m examples.scan_data
License
MIT License
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file filescan-0.0.3.tar.gz.
File metadata
- Download URL: filescan-0.0.3.tar.gz
- Upload date:
- Size: 12.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
55f8a996176c0b95fb154190eec7e6d672e24e5814088afae425caeb73e937d2
|
|
| MD5 |
7f9a5853c794e69c06980a58d97594a0
|
|
| BLAKE2b-256 |
6978b9049e1862b6bc44a9f8056888cb34fdb4303afb9bd0a799038232e0b3f5
|
Provenance
The following attestation bundles were made for filescan-0.0.3.tar.gz:
Publisher:
publish.yml on DreamSoul-AI/filescan
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
filescan-0.0.3.tar.gz -
Subject digest:
55f8a996176c0b95fb154190eec7e6d672e24e5814088afae425caeb73e937d2 - Sigstore transparency entry: 945031345
- Sigstore integration time:
-
Permalink:
DreamSoul-AI/filescan@5b38bb559050ac412b4b60c6002168954f034781 -
Branch / Tag:
refs/tags/v0.0.3 - Owner: https://github.com/DreamSoul-AI
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@5b38bb559050ac412b4b60c6002168954f034781 -
Trigger Event:
push
-
Statement type:
File details
Details for the file filescan-0.0.3-py3-none-any.whl.
File metadata
- Download URL: filescan-0.0.3-py3-none-any.whl
- Upload date:
- Size: 11.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0a9f8707167346cdb08bb632f10c895585f4dbafc1be685d1a357e5f9949da9b
|
|
| MD5 |
f87c96f90ebe792eaf76a2032fa4fedb
|
|
| BLAKE2b-256 |
50b081274c2bda7fe2722d10c2788f74140a3e5d5f3c8c9ca2a9815242a82015
|
Provenance
The following attestation bundles were made for filescan-0.0.3-py3-none-any.whl:
Publisher:
publish.yml on DreamSoul-AI/filescan
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
filescan-0.0.3-py3-none-any.whl -
Subject digest:
0a9f8707167346cdb08bb632f10c895585f4dbafc1be685d1a357e5f9949da9b - Sigstore transparency entry: 945031412
- Sigstore integration time:
-
Permalink:
DreamSoul-AI/filescan@5b38bb559050ac412b4b60c6002168954f034781 -
Branch / Tag:
refs/tags/v0.0.3 - Owner: https://github.com/DreamSoul-AI
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@5b38bb559050ac412b4b60c6002168954f034781 -
Trigger Event:
push
-
Statement type: