Skip to main content

Simple minded facilities for media information inferred from filenames. This contains mostly lexical functions for extracting information from strings or constructing media filenames from metadata and a few classes like EpisodeInfo and SeriesEpisodeInfo for common descriptions.

The default filename parsing rules are based on my personal convention, which is to name media files as:

series_name--episode_info--title--source--etc-etc.ext

where the components are:

  • series_name: the programme series name downcased and with whitespace replaced by dashes; in the case of standalone items like movies this is often the studio.
  • episode_info: a structured field with episode information: sn is a series/season, enis an episode number within the season,x_n_ is a "extra" - addition material supplied with the season, etc.
  • title: the episode title downcased and with whitespace replaced by dashes.
  • source: the source of the media.
  • ext: filename extension such as mp4.

As you may imagine, as a rule I dislike mixed case filenames and filenames with embedded whitespace. I also like a media filename to contain enough information to identify the file contents in a compact and human readable form.

Short summary:

  • EpisodeDatumDefn: An EpisodeInfo marker definition with the following components: - name: the marker name, such as "series" or "episode" - prefix: the stub used in a filename, such as "s" or "e" - re: a regular expression to match the prefix an some digits.

  • EpisodeInfo: Trite class for episodic information, used to store, match or transcribe series/season, episode, etc values.

  • main: Main command line running some test code.

  • parse_name: Parse the descriptive part of a filename (the portion remaining after stripping the file extension) and yield (part,fields) for each part as delineated by sep.

  • part_to_title: Convert a filename part into a title string.

  • pathname_info: Parse information from the basename of a file pathname. Return a mapping of field => values in the order parsed.

  • scrub_title: Strip redundant text from the start of an episode title.

  • SeriesEpisodeInfo: Episode information from a TV series episode.

  • title_to_part: Convert a title string into a filename part. This is lossy; the part_to_title function cannot completely reverse this.

Functions

main(argv=None)

Main command line running some test code.

parse_name(name, sep='--')

Parse the descriptive part of a filename (the portion remaining after stripping the file extension) and yield (part,fields) for each part as delineated by sep.

part_to_title(part)

Convert a filename part into a title string.

Example:

>>> part_to_title('episode-name')
'Episode Name'

pathname_info(pathname)

Parse information from the basename of a file pathname. Return a mapping of field => values in the order parsed.

scrub_title(title: str, *, season=None, episode=None) -> str

Strip redundant text from the start of an episode title.

I frequently get "title" strings with leading season/episode information. This function cleans up these strings to return the unadorned title.

title_to_part(title)

Convert a title string into a filename part. This is lossy; the part_to_title function cannot completely reverse this.

Example:

>>> title_to_part('Episode Name')
'episode-name'

Classes

class EpisodeDatumDefn(EpisodeDatumDefn)

An EpisodeInfo marker definition with the following components:

  • name: the marker name, such as "series" or "episode"
  • prefix: the stub used in a filename, such as "s" or "e"
  • re: a regular expression to match the prefix an some digits

EpisodeDatumDefn.parse(self, s, offset=0)

Parse an episode datum from a string, return the value and new offset. Raise ValueError if the string doesn't match this definition.

Parameters:

  • s: the string
  • offset: parse offset, default 0

class EpisodeInfo(types.SimpleNamespace)

Trite class for episodic information, used to store, match or transcribe series/season, episode, etc values.

EpisodeInfo.MARKERS

[   EpisodeDatumDefn(name='series', prefix='s', re=re.compile('s(\\d+)', re.IGNORECASE)),
    EpisodeDatumDefn(name='episode', prefix='e', re=re.compile('e(\\d+)', re.IGNORECASE)),
    EpisodeDatumDefn(name='part', prefix='pt', re=re.compile('pt(\\d+)', re.IGNORECASE)),
    EpisodeDatumDefn(name='part', prefix='p', re=re.compile('p(\\d+)', re.IGNORECASE)),
    EpisodeDatumDefn(name='scene', prefix='sc', re=re.compile('sc(\\d+)', re.IGNORECASE)),
    EpisodeDatumDefn(name='extra', prefix='x', re=re.compile('x(\\d+)', re.IGNORECASE)),
    EpisodeDatumDefn(name='extra', prefix='ex', re=re.compile('ex(\\d+)', re.IGNORECASE))]

EpisodeInfo.__getitem__(self, name)

We can look up values by name.

EpisodeInfo.as_dict(self)

Return the episode info as a dict.

EpisodeInfo.as_tags(self, prefix=None)

Generator yielding the episode info as Tags.

EpisodeInfo.from_filename_part(s, offset=0)

Factory to return an EpisodeInfo from a filename episode field.

Parameters:

  • s: the string containing the episode information
  • offset: the start of the episode information, default 0

The episode information must extend to the end of the string because the factory returns just the information. See the parse_filename_part class method for the core parse.

EpisodeInfo.get(self, name, default=None)

Look up value by name with default.

EpisodeInfo.parse_filename_part(s, offset=0)

Parse episode information from a string, returning the matched fields and the new offset.

Parameters: s: the string containing the episode information. offset: the starting offset of the information, default 0.

EpisodeInfo.season

<property object at 0x105e99e90>

class SeriesEpisodeInfo(cs.deco.Promotable)

Episode information from a TV series episode.

SeriesEpisodeInfo.__hash__

None

SeriesEpisodeInfo.as_dict(self)

Return the non-None values as a dict. Note that this uses dataclasses.asdict() and as such is a deep copy.

SeriesEpisodeInfo.episode_part

None

SeriesEpisodeInfo.episode_title

None

SeriesEpisodeInfo.from_str(episode_title: str, series=None)

Infer a SeriesEpisodeInfo from an episode title.

This recognises the common 'sSSeEE - Episode Title' format and variants like Series Name - sSSeEE - Episode Title' or 'sSSeEE - Episode Title - Part: One'.

SeriesEpisodeInfo.season_title

None

Release files for cs-mediainfo 20260914

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cs-mediainfo 20260914
File Size Uploaded
cs_mediainfo-20260914.tar.gz 6.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cs-mediainfo 20260914
File Interpreter ABI Platform
cs_mediainfo-20260914-py2.py3-none-any.whl Python 3, Python 2 none any Details

Total release size: 14.8 kB

Release files / cs_mediainfo-20260914.tar.gz

Download URL cs_mediainfo-20260914.tar.gz
Size 6.6 kB
Tags Source
SHA-256 checksum
How to use checksums
cbe42af0225caae77c82977935556fa0d62b4f80da2e73fadb2047f97a5aa1c7
BLAKE2b-256 checksum
How to use checksums
3dedb3537cf26212a9081faf461d845f4e88f99597c54d5f8b7595b1d7c2b4b7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.1

Release files / cs_mediainfo-20260914-py2.py3-none-any.whl

Download URL cs_mediainfo-20260914-py2.py3-none-any.whl
Size 8.2 kB
Tags Python 2 Python 3
SHA-256 checksum
How to use checksums
a723c51513e9fb5b89ecf6958945df240e8f4361c9e5fd24e1af3f4a23f50ef1
BLAKE2b-256 checksum
How to use checksums
acafa0ed9d7fe79ae0171858a68a027845b8cb8606221371729ee9aa30cf4f6e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.13.1

Release history Release notifications | RSS feed

This release

20260914 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page