Skip to main content

Secure and extensible directory extraction pipeline with filtering and guardrails

Project description

📂 Directory Extraction Mechanism Using Python

A robust, secure, and flexible directory scanning pipeline built with FastAPI that extracts files from a directory using multiple filtering strategies.

It supports:

  • ✅ Multiple selection modes (all / by types / by names / by patterns)
  • ✅ Exclusion filters
  • ✅ Safety guardrails (size, date range, result limit)
  • ✅ Secure path validation (prevents directory traversal)
  • ✅ Clean structured JSON response with detailed stats

🚀 Features

🔎 1. Multiple Selection Modes

You can extract files using different strategies:

Mode Description
all Return all files recursively
by_types Filter by file extensions (pdf, md, docx, etc.)
by_names Select exact filenames or relative paths
by_patterns Use glob patterns like **/*.md

🛡 2. Built-in Guardrails

After selection and exclusion, files are validated against:

  • 📏 Maximum file size (max_file_size_mb)
  • 📅 Modification time window (modified_after, modified_before)
  • 🔢 Maximum number of results (limit)
  • ❌ Files deleted during processing (tracked safely)

🔐 3. Security First

  • Uses Path.resolve() to avoid path traversal attacks
  • Rejects absolute glob patterns
  • Ensures all files remain inside the target directory
  • Avoids symlink loops

🏗 Project Structure

.
├── fastapi_main_app.py   # API entrypoint and pipeline orchestration
└── business_logic
    ├── schemas.py        # Pydantic models & enums
    ├── utils.py          # Core file collection & filtering logic

⚙️ How the Pipeline Works

The API follows a 4-step processing pipeline :

1️⃣ SELECT files (based on mode)
2️⃣ APPLY EXCLUSIONS
3️⃣ APPLY GUARDRAILS
4️⃣ FORMAT RESPONSE

Step 1 — Selection

Depending on mode:

  • all → Collect everything recursively
  • by_types → Filter by extensions
  • by_names → Match exact filenames or relative paths
  • by_patterns → Match glob patterns

Step 2 — Exclusions

Optional exclude_names supports:

  • Relative paths → docs/README.md
  • Bare filenames → temp.txt (removes everywhere)
  • Case-insensitive matching (optional)

Step 3 — Guardrails

Applies:

  • File size limit
  • Date filtering (UTC normalized)
  • Result cap
  • Tracks skipped files with reasons

Step 4 — Response Formatting

Returns:

  • Total candidates
  • Exclusion stats
  • Guardrail skip stats
  • Final selected files (relative + absolute paths)

📦 Installation

git clone https://github.com/siddharth1310/directory_extraction_mechanism
cd https://github.com/siddharth1310/directory_extraction_mechanism
python -m venv venv
source venv/bin/activate  # Windows: venv\Scripts\activate
pip install -r requirements.txt

Run the server:

uvicorn fastapi_main_app:app --reload

Open docs:

http://127.0.0.1:8000/docs

🧠 API Usage

Endpoint

POST /directory/

Example Request

{
  "directory": "/home/user/docs",
  "mode": "by_types",
  "file_types": ["pdf", "md"],
  "max_file_size_mb": 10,
  "limit": 50
}

Example Response

{
  "total_candidates": 127,
  "excluded_names": [],
  "excluded_names_matched": [],
  "post_exclude_candidates": 115,
  "total_selected": 42,
  "selected_files": [
    "README.md",
    "docs/guide.pdf"
  ],
  "skipped": {
    "too_large": 5,
    "outside_time_window": 68,
    "disappeared": 0,
    "files": []
  },
  "selected_files_path": [
    "/home/user/docs/README.md",
    "/home/user/docs/docs/guide.pdf"
  ]
}

📘 Request Model Explained

Required

Field Type Description
directory string Root directory to scan

Selection Options

Field Used When Description
mode Always Selection strategy
file_types by_types Extensions to include
file_names by_names Exact names or paths
include_globs by_patterns Glob patterns

Exclusions

Field Description
exclude_names Names or relative paths to remove
case_insensitive Only for by_names mode
duplicate_policy error,first,all

Guardrails

Field Description
max_file_size_mb Skip large files
modified_after Include files newer than
modified_before Include files older than
limit Maximum number of files

🧩 Duplicate Handling (by_names mode)

When multiple files share the same name:

  • error → Throw validation error
  • first → Return first match
  • all → Return all matches (default)

🌍 Glob Pattern Examples

Pattern Meaning
**/*.md All markdown files recursively
docs/**/*.pdf PDFs inside docs folder
*.txt Text files in root only

🧪 Example Use Cases

✅ Index all files

{
  "directory": "/data"
}

✅ Only PDFs modified after Jan 1, 2024

{
  "directory": "/data",
  "mode": "by_types",
  "file_types": ["pdf"],
  "modified_after": "2024-01-01T00:00:00"
}

✅ Extract specific files safely

{
  "directory": "/project",
  "mode": "by_names",
  "file_names": ["README.md", "docs/guide.pdf"],
  "duplicate_policy": "first"
}

📊 Why This Project is Robust

  • Memory-efficient generator-based scanning
  • Safe path containment validation
  • Clean separation of concerns
  • Clear stats for observability
  • Threadpool execution for blocking file I/O
  • Fully validated request model via Pydantic

🔮 Future Enhancements (Ideas)

  • Async file scanning
  • Streaming large directory results
  • Caching indexed directories
  • File hashing support
  • Logging & observability hooks

🤝 Contributing

Contributions, issues, and feature requests are welcome! Feel free to fork the repo and submit a pull request.


🧾 License

Use it freely for research and development.


👤 Author

Created by Siddharth Singh. Find me on LinkedIn

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

directory_extractor-0.1.0.tar.gz (16.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

directory_extractor-0.1.0-py3-none-any.whl (16.5 kB view details)

Uploaded Python 3

File details

Details for the file directory_extractor-0.1.0.tar.gz.

File metadata

  • Download URL: directory_extractor-0.1.0.tar.gz
  • Upload date:
  • Size: 16.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for directory_extractor-0.1.0.tar.gz
Algorithm Hash digest
SHA256 178a0a6465cdae81a9e3fcbbf193657460965e71421f9778b4d35b4f377c3ad5
MD5 f031b41903bfc11cd57fc64192109580
BLAKE2b-256 77c047eae867fb65d9bf477047940013fa29288ac9686b8cbd8f89f31a1c5fd8

See more details on using hashes here.

File details

Details for the file directory_extractor-0.1.0-py3-none-any.whl.

File metadata

File hashes

Hashes for directory_extractor-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 f88d3c817b11f059961d028f07a457f5d9bd9e4d1579965fce59bb8eb1ea58ed
MD5 35aaecb2a1f20f2d571f2fbbe61e0bc5
BLAKE2b-256 3d194b7daf1a73cd59a38c3c83459b7e2376d62a69f8f9126a747983dad93357

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page