Skip to main content

A tool to extract content from text files in a directory, with options to ignore certain files and directories.

Project description

FileContentExtractor

Languages: 🇺🇸English | 🇰🇷한국어

FileContentExtractor is a versatile Python tool designed to recursively extract and save the contents of text files in a directory. It provides various options to ignore specific files or directories, filter files based on several criteria, and output file paths as either absolute or relative paths. The tool also includes user confirmation, version display, and help options for better usability.

Features

  • Recursively traverses directories and extracts the content of text files.
  • Allows you to ignore specific files, directories, or patterns using a .ignorelist file.
  • Outputs file paths as either relative or absolute paths.
  • Supports filtering files by extension, size, and modification date.
  • Allows splitting output into multiple files based on size.
  • Displays a progress bar during extraction.
  • Prompts user confirmation before proceeding with extraction.
  • Provides version information and help details via command-line options.

Installation

You can install the package via pip:

pip install file-content-extractor

Usage

  1. Basic command:

    • Default usage (relative paths):

      file_extractor
      
      • After running this command, you will be prompted to confirm the extraction by typing y.
    • Use absolute paths:

      file_extractor -a
      
  2. Specify additional options:

    • Ignore specific files or directories:

      file_extractor -i my_ignore_list.txt
      
    • Save output to a custom file:

      file_extractor -o my_output.txt
      
    • Filter by file extension, size, and modification date:

      file_extractor -e .txt,.py -m 1024 -M 1048576 -d 2023-01-01
      
    • Split output into multiple files:

      file_extractor -s 10485760
      
    • Specify the directory to start extraction from:

      file_extractor -p /path/to/directory
      
  3. Check version:

    • To display the current version of the tool:

      file_extractor -v
      
  4. Help:

    • To display help information and usage details:

      file_extractor -h
      

Options

  • -a, --absolute: Use absolute paths in the output.
  • -i <file>, --ignore=<file>: Specify a custom ignore file (default is .ignorelist).
  • -o <file>, --output=<file>: Specify a custom output file (default is output.txt).
  • -e <ext1,ext2,...>, --extensions=<ext1,ext2,...>: Specify file extensions to include (e.g., .txt,.py).
  • -m <bytes>, --min-size=<bytes>: Specify minimum file size to include.
  • -M <bytes>, --max-size=<bytes>: Specify maximum file size to include.
  • -d <YYYY-MM-DD>, --modified-after=<YYYY-MM-DD>: Include only files modified after a specific date.
  • -s <bytes>, --split-size=<bytes>: Split output files into chunks of the specified size.
  • -p <directory>, --path=<directory>: Specify the directory to start extraction from.
  • -h, --help: Show this help message and exit.
  • -v, --version: Show version information and exit.

Contributing

Contributions are welcome! Feel free to submit a Pull Request.

License

This project is licensed under the MIT License. See the LICENSE file for details.

Author

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

file_content_extractor-0.2.4.tar.gz (7.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

file_content_extractor-0.2.4-py3-none-any.whl (7.8 kB view details)

Uploaded Python 3

File details

Details for the file file_content_extractor-0.2.4.tar.gz.

File metadata

  • Download URL: file_content_extractor-0.2.4.tar.gz
  • Upload date:
  • Size: 7.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.12.4

File hashes

Hashes for file_content_extractor-0.2.4.tar.gz
Algorithm Hash digest
SHA256 687fe02e5dba58f2de66dcfd5da465e6cfd38c6e0c08a0adef0424430c02da29
MD5 1fbbd6e6c9c5ceec9dd2dd4e0a47da1b
BLAKE2b-256 8ebcd28ea5364695eacd73ad6e3ffb06d503d2ba528d93f740069f0334d21e45

See more details on using hashes here.

File details

Details for the file file_content_extractor-0.2.4-py3-none-any.whl.

File metadata

File hashes

Hashes for file_content_extractor-0.2.4-py3-none-any.whl
Algorithm Hash digest
SHA256 03a9cde49f0945fef7f8c6f5020f8f77a97d8c95076ab1eff4f0b8630ee19238
MD5 50cae503cd5adef87d09f14b8fff0902
BLAKE2b-256 1c4f05a1cc9742c0432c9402222d75b702e5a5dd9a3cc441c1095ef068be3190

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page