Skip to main content

Composable neural network components for building models in PyTorch.

Project description

composennent

PyPI version Python 3.8+ License: MIT

Composable neural network components for building models in PyTorch.

Composennent provides modular, reusable building blocks for constructing transformer-based models. Train GPT, BERT, and other architectures with minimal code.

Features

  • 🧩 Modular Components: Encoder, Decoder, Attention blocks that compose together
  • 🚀 Built-in Training: Pre-training and fine-tuning with a single method call
  • 📝 Multiple Architectures: GPT, BERT, Seq2Seq support out of the box
  • 🔧 Tokenizer Support: WordPiece and SentencePiece tokenizers included
  • Mixed Precision: Automatic mixed precision (AMP) support
  • 🎯 Instruction Tuning: Fine-tune models on instruction datasets (Alpaca format)

Installation

pip install composennent

For tokenizer support:

pip install composennent[tokenizers]

For development:

pip install composennent[dev]

Quick Start

Pre-train a GPT Model

import torch
from composennent.nlp.transformers import GPT
from composennent.nlp.tokenizers import SentencePieceTokenizer

# Create model
model = GPT(
    vocab_size=32000,
    latent_dim=512,
    num_heads=8,
    num_layers=6,
    max_seq_len=512,
)

# Load tokenizer
tokenizer = SentencePieceTokenizer.from_pretrained("tokenizer.model")

# Pre-train
texts = ["Your training data here...", ...]
model.pretrain(
    texts=texts,
    tokenizer=tokenizer,
    epochs=3,
    batch_size=16,
    device="cuda",
)

# Save
model.save("my_model.pt")

Fine-tune on Instructions

# Load pre-trained model
model = GPT.load("my_model.pt", device="cuda")

# Instruction data (Alpaca format)
instruction_data = [
    {
        "instruction": "What is the capital of France?",
        "input": "",
        "output": "The capital of France is Paris."
    },
    # ... more examples
]

# Fine-tune
model.fine_tune(
    data=instruction_data,
    tokenizer=tokenizer,
    epochs=2,
    lr=5e-5,
    mask_prompt=True,  # Only compute loss on outputs
)

Generate Text

prompt = tokenizer.encode("What is")
generated = model.generate(
    input_ids=prompt,
    max_length=100,
    temperature=0.8,
)
print(tokenizer.decode(generated[0].tolist()))

Modules

Module Description
composennent.basic Core building blocks (Encoder, Decoder, Block)
composennent.attention Attention mechanisms and masks
composennent.nlp.transformers GPT, BERT, and other transformer models
composennent.nlp.tokenizers WordPiece and SentencePiece tokenizers
composennent.training Training utilities and trainer classes
composennent.expert Mixture of Experts components
composennent.vision Vision transformer components
composennent.utils Utility functions

Training API

For more control over training, use the trainer classes directly:

from composennent.training import CausalLMTrainer, train

# Option 1: Use the train() convenience function
train(model, texts, tokenizer, model_type="causal_lm", epochs=5)

# Option 2: Use trainer class directly
trainer = CausalLMTrainer(model, tokenizer, device="cuda")
trainer.train(texts, epochs=5, batch_size=16)
trainer.save_checkpoint("checkpoint.pt")

Available trainers:

  • CausalLMTrainer - GPT-style next-token prediction
  • MaskedLMTrainer - BERT-style masked language modeling
  • Seq2SeqTrainer - Encoder-decoder models
  • MultiTaskTrainer - Multi-task learning (MLM + NSP)
  • CustomTrainer - Custom loss functions

Requirements

  • Python >= 3.8
  • PyTorch >= 2.0.0

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/amazing-feature)
  3. Install dev dependencies (pip install -e ".[dev]")
  4. Run tests (pytest)
  5. Run formatters (black . && ruff check .)
  6. Commit your changes (git commit -m 'Add amazing feature')
  7. Push to the branch (git push origin feature/amazing-feature)
  8. Open a Pull Request

License

MIT License - see LICENSE for details.

Links

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

composennent-0.4.0.tar.gz (53.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

composennent-0.4.0-py3-none-any.whl (70.1 kB view details)

Uploaded Python 3

File details

Details for the file composennent-0.4.0.tar.gz.

File metadata

  • Download URL: composennent-0.4.0.tar.gz
  • Upload date:
  • Size: 53.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for composennent-0.4.0.tar.gz
Algorithm Hash digest
SHA256 2e40b3bdac72cbf3560fe7890b8e8f98064a7d5f8439e7e6a609215b74130859
MD5 12cbf6b11404d98d520aa2043494466a
BLAKE2b-256 10505a78ec7528aa3ba5fc6e79c66cce5f21b56c5857730ef782e86db72292e0

See more details on using hashes here.

File details

Details for the file composennent-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: composennent-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 70.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for composennent-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 4f0d8fca84eb7c48d5a9c695427015979a2e612a95d8388983a241492a113a30
MD5 9efe8ec9c8f3dc88645537c1497c3d88
BLAKE2b-256 dd0a60a45e2b3a577ba780921696cf081b6d0589aa63c00dd2a3105c9794361b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page