Skip to main content

Meaning-Informed Next-token Transformation

Project description

Publish to PyPI CI

MINT

Meaning-Informed Next-token Transformation

Project Goals

MINT adds a transformation layer that redistributes next-token probabilities according to semantic token similarity based on the model's embedding space. The aim is to produce more varied, human-like text without sacrificing coherence.

Installation

Create a virtual environment and install the package in editable mode:

pip install -e .

This installs the mint package and provides the mint command-line interface.

To use the published release from PyPI (when available), run:

pip install mint-llm

CLI Usage

Run the CLI with:

mint --help

The mint command exposes several subcommands. The typical workflow is shown below.

Using MINT

Run the commands below to build and apply the redistribution layer:

  1. Pick a checkpoint from the Hugging Face Hub (optional).

    mint pick <model_id> checkpoint/
    
  2. Extract token embeddings from the checkpoint.

    mint extract checkpoint/model.safetensors embeddings.safetensors
    
  3. Blend the embeddings into a low-rank similarity factor.

    mint blend embeddings.safetensors mint_out/ --rank 1024
    # → mint_out/W.safetensors (plus R.safetensors with --keep-residual)
    

    Use -r/--rank to set factor rank (default 1024). Pass --keep-residual to also save a sparse R.safetensors file.

  4. Brew new text from the wrapped model.

    mint brew model_id_or_path mint_out/ --prompt "Hello"
    

    Omit --prompt or pass --interactive to read prompts from stdin.

  5. Infuse the tested similarity matrix into a local model and save the result to a directory.

    mint infuse path/to/model mint_out/ infused-model --alpha 0.1
    
    from mint.wrapper import load_wrapped_model
    from mint.logits import SRLogitsProcessor
    
    model, tokenizer, layer = load_wrapped_model("model_id_or_path", "mint_out/")
    processor = SRLogitsProcessor(layer)
    

See the notebooks and examples/quickstart.py for a more detailed walk-through and an automated script. You can also explore the generator interactively using the CLI.

Additional Utilities

The CLI exposes optional commands for working with checkpoints:

  • Crush merge sharded checkpoints referenced by an index file.

    mint crush checkpoint/model.safetensors.index.json checkpoint/model.safetensors
    
  • Chop split a .safetensors checkpoint into shards. Provide a shard count or size:

    mint chop model.safetensors shards/ --shards 2
    
    mint chop model.safetensors shards/ --size-mb 500
    

Quickstart Script

Run examples/quickstart.py for an end-to-end demonstration. The script mirrors the mint CLI commands: extract, blend and brew.

Required argument:

  • --prompt – input text to generate from.

Optional arguments default to values defined in tests/utils/model_config.json:

  • --checkpoint – checkpoint path. If this points to a *.safetensors.index.json file the required shards are downloaded and merged automatically. If omitted model_url is used.
  • --model – model identifier or path. When a model ID is provided the checkpoint shards are fetched and merged automatically. Defaults to model_id or one derived from model_url.
  • --embeddings – output file for embeddings (default embeddings.safetensors).
  • --similarity – output directory for W.safetensors (and optionally R.safetensors, default .cache/mint).
python examples/quickstart.py --prompt "Hello"

The script extracts embeddings, builds the similarity matrix and generates text using the wrapped model.

Examples

Practical examples are provided in the notebooks directory. They demonstrate embedding extraction, building a similarity matrix and brewing text from a short prompt.

Development

Install development dependencies with:

pip install -e '.[dev]'

Use the provided Makefile to run common tasks:

make format # check black formatting
make lint   # run ruff and mypy (if configured)
make test   # run the pytest suite
make all    # runs all checks

make commands format, lint, and all can also be suffixed with -fix (e.g. make format-fix) to attempt to automatically fix issues. make fix will also run all fixes.

Contributing

Development tasks are tracked in todos.json. See project_proposal-MINT.md for the full technical plan. Release notes are available in CHANGELOG.md. Feel free to open issues or pull requests to contribute.

Citation

If you use MINT in your research, please cite the project using the metadata in CITATION.cff.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mint_llm-0.0.3.tar.gz (20.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mint_llm-0.0.3-py3-none-any.whl (17.7 kB view details)

Uploaded Python 3

File details

Details for the file mint_llm-0.0.3.tar.gz.

File metadata

  • Download URL: mint_llm-0.0.3.tar.gz
  • Upload date:
  • Size: 20.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for mint_llm-0.0.3.tar.gz
Algorithm Hash digest
SHA256 8db86fc623f67b34d6eda53b4b8669b0610aa2adcf1329305954a359ff8d8131
MD5 c43b1ca7eaff2f5b27c66fdee17015b8
BLAKE2b-256 cb7b2868b5ed54e4a08b0bff77b33eed6364c17fdd4608f7f7c6edfb37a22898

See more details on using hashes here.

Provenance

The following attestation bundles were made for mint_llm-0.0.3.tar.gz:

Publisher: publish.yml on Reithan/MINT

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mint_llm-0.0.3-py3-none-any.whl.

File metadata

  • Download URL: mint_llm-0.0.3-py3-none-any.whl
  • Upload date:
  • Size: 17.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for mint_llm-0.0.3-py3-none-any.whl
Algorithm Hash digest
SHA256 bc6d70225ec8e79e9763b807057f6a2c3e535328fad9b3060b724418c738cc6f
MD5 3e8293c9672f846064e17d9635234139
BLAKE2b-256 e0ff691e903a052e4fd88b58ee4402954eb02377d6d4367f05c19580af24c092

See more details on using hashes here.

Provenance

The following attestation bundles were made for mint_llm-0.0.3-py3-none-any.whl:

Publisher: publish.yml on Reithan/MINT

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page