Skip to main content

Subtitles extremely clean

Project description

Subtitles extremely clean.

Latest Version Travis CI build status License
Project page:

https://github.com/ratoaq2/cleanit

CleanIt is a command line tool (written in python) that helps you to keep your subtitles clean. You can specify rules to detect subtitle entries to be removed or patterns to be replaced. Simple text matching or complex regex can be used.

Usage

CLI

Clean subtitles:

$ cleanit --config my-config.yml my-subtitle.srt
Collected 1 subtitles
Saving <Subtitle [my-subtitle.srt]>
Saved <Subtitle [my-subtitle.srt]>

Library

How to clean subtitles in a specific path using a specific configuration:

from cleanit.api import clean_subtitle, save_subtitle
from cleanit.config import Config
from cleanit.subtitle import Subtitle

subtitle = Subtitle('/subtitle/path')
config = Config.from_file('/config/path')
if clean_subtitle(subtitle, config.rules):
    save_subtitle(subtitle)

YAML Configuration file

The yaml configuration file has 2 main sections: templates and groups.

  • Templates can help you to define common configuration snippets to be used in several groups.

  • Groups: where you can define your rules.

# Reference:
#   type: [text*, regex]
#   match: [contains*, exact, startswith, endswith]
#   flags: [ignorecase, dotall, multiline, locale, unicode, verbose]
#   whitelist: no*
#   rules:
#   - sometext
#   - (\b)(\d{1,2})x(\d{1,2})(\b): {replacement: \1S\2E\3\4, type: regex, match: contains, flags: [unicode], whitelist: no}


templates:
  common:
    type: text
    match: contains

groups:
  # Groups can have any name, in this case 'blacklist' we have all the rules to remove subtitle  entries
  blacklist:
    template: common
    rules:
      # Removes any subtitle entry that contains the word FooBar
      - FooBar

      # Removes any subtitle entry that contains the pattern S00E00
      # Example:
      #   My Series S01E02
      - \bs\d{2}\s?e\d{2}\b: {type: regex, flags: ignorecase}

      # Removes any subtitle entry that is exactly the word: 'Ah' or 'Oh' (with 1 or more h)
      # Example:
      #   Ohhh!
      - ((Ah+)|(Oh+))\W?: {match: exact}

  # The group 'tidy' has all rules to replace certain patterns in your subtitles.
  tidy:
    template: common
    type: regex
    rules:
      # Description: Replace extra spaces to a single space
      # Example:
      #   Foo     bar.
      # to
      #   Foo bar.
      - \s{2,}: ' '

      # Description: Add space when starting phrase with '-'. It ignores tags, such as <i>, <b>
      # Example:
      #   <i>-Francine, what has happened?
      #   -What has happened? You tell me!</i>
      # to
      #   <i>- Francine, what has happened?
      #   - What has happened? You tell me!</i>
      - '(?:^(|(?:\<\w\>)))-([''"]?\w+)': { replacement: '\1- \2', flags: [multiline, unicode] }

* The default value if none is defined

CleanIt will try to load configuration file from ~/.config/cleanit/config.yml if no configuration file is defined.

Changelog

0.2.1

release date: 2016-02-28 * Adding guess encoding back without python-magic dependency.

0.2

release date: 2016-02-27 * Removing chardet and python-magic dependencies. Either encoding is specified or it should be guessed by pysrt

0.1

release date: 2015-10-16

  • Initial release

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cleanit-0.2.1.tar.gz (11.6 kB view details)

Uploaded Source

File details

Details for the file cleanit-0.2.1.tar.gz.

File metadata

  • Download URL: cleanit-0.2.1.tar.gz
  • Upload date:
  • Size: 11.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No

File hashes

Hashes for cleanit-0.2.1.tar.gz
Algorithm Hash digest
SHA256 6a7d05fbaa1506235036514edd6ceabe7e407de5557269bbd7162d0a8bd95086
MD5 0f109cb3e1399ee078c723c8ce7007fb
BLAKE2b-256 01b962cadbaa776e83b722c7ac7d53341107df36a771ef293b28934c5fd01094

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page