Add your description here
Project description
AMALIA
AMALIA is a decoder-only transformer architecture inherited from EuroLLM-9B, implemented in plain PyTorch (no external attention kernels).
Local usage (with uv)
Install uv:
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
Clone the repo and run the example:
git clone https://github.com/tiagomonteiro0715/amalia-core.git
cd amalia-core
uv sync
uv run main.py
uv sync installs the dependencies from pyproject.toml/uv.lock into a local .venv, and uv run executes inside it without needing to activate it manually.
Usage in Google Colab
!pip install uv
!uv pip install --system amalia
import torch
from amalia import AmaliaConfig, AmaliaForCausalLM
# Initialize the model with random weights
config = AmaliaConfig()
model = AmaliaForCausalLM(config)
# Run a forward pass on random token ids
input_ids = torch.randint(0, config.vocab_size, (1, 16))
logits = model(input_ids)
print(logits.shape) # torch.Size([1, 16, 128000])
Project structure
amalia/
├── amalia/ # the installable library
│ ├── __init__.py # public API surface
│ ├── config.py # hyperparameters
│ └── architecture.py # the model itself
├── main.py # runnable usage example
├── pyproject.toml # package metadata + build config
├── release.bat # version bump / GitHub tag / PyPI publish script
└── LICENSE
The library is split into a config module and an architecture module rather than one file because the two change for different reasons: hyperparameters get tuned/swapped constantly, while the model code (the actual math) is comparatively stable. Keeping them separate means you can hand AmaliaConfig around (e.g. save it as JSON alongside a checkpoint) without dragging PyTorch nn.Module code with it.
amalia/config.py
AmaliaConfig— a plain@dataclassholding every architecture hyperparameter (vocab size, hidden size, number of layers/heads, RoPE θ, etc.), with sensible defaults matching the AMALIA spec. It's a dataclass instead of a bigger config framework (like HuggingFace'sPretrainedConfig) because there's no need for serialization/versioning machinery yet — a dataclass gives typed fields, a free__init__, and a readablerepr()for zero extra code.head_dimis a@property(hidden_size // num_attention_heads) instead of a stored field, because it's fully determined by the other two — storing it separately would let it silently drift out of sync if someone editedhidden_sizeafter construction.
amalia/architecture.py
Built bottom-up: small standalone pieces first, composed into progressively bigger nn.Modules. Each class exists because it's a distinct, reusable computation with its own shape/semantics — splitting them keeps each one testable and readable in isolation instead of one monolithic forward.
RMSNorm— root-mean-square layer normalization (like LayerNorm but without re-centering on the mean). It's its own class because it's used twice per decoder layer (before attention, before the MLP) plus once at the very end of the model — defining it once avoids repeating the same three lines everywhere.RotaryEmbedding— precomputes the inverse-frequency table used by RoPE (Rotary Position Embeddings) fromhead_dimandrope_theta, then producescos/sintensors for a given sequence length. It's a module (not a plain function) soinv_freqcan live as a registered buffer — it moves with the model when you call.to(device), but isn't a trainable parameter.rotate_half/apply_rotary_pos_emb— free functions, not methods, because they're pure tensor math with no state, applied identically to both queries and keys. Keeping them as standalone functions (rather than duplicating the logic insideGroupedQueryAttention) makes the RoPE math easy to unit-test on its own.repeat_kv— expands the key/value heads so they broadcast against the (larger) number of query heads. It's a separate function because Grouped Query Attention (GQA) is the one part of the attention block that isn't "standard" multi-head attention, so isolating it makes the GQA-specific logic obvious at a glance.GroupedQueryAttention— the attention block: projectsxinto Q/K/V (Q has more heads than K/V, per GQA), applies RoPE, repeats K/V to match Q's head count, then calls PyTorch's built-inF.scaled_dot_product_attentionwithis_causal=True. This is the "no external flash-attention library" requirement in practice —scaled_dot_product_attentionships in core PyTorch and picks an efficient fused kernel under the hood, without adding a dependency likeflash-attn.SwiGLU— the feed-forward block:down_proj(silu(gate_proj(x)) * up_proj(x)), the gated activation used by Llama-family models (in place of a plain ReLU/GELU MLP) because it consistently improves quality for the same parameter budget.DecoderLayer— one transformer block: pre-norm residual attention, then pre-norm residual MLP. "Pre-norm" (normalize before the sub-layer, not after) is what makes it feasible to stack 42 of these without training instability.AmaliaModel— the backbone: token embedding → 42 stackedDecoderLayers → finalRMSNorm. It stops at hidden states and deliberately has no output head, so the same backbone could in principle back other heads later (e.g. a classification head) without change.AmaliaForCausalLM— wrapsAmaliaModelwith anlm_head(aLinearprojecting hidden states to vocabulary logits) and is the class you actually instantiate for language modeling. The head is a separateLinear, not tied to the embedding weights, because the spec calls for untied embeddings (tie_word_embeddings = False).
The forward pass, end to end (AmaliaForCausalLM.forward): embed token ids → compute RoPE cos/sin once for the sequence length → run through all 42 decoder layers → final norm → project to vocab logits. There's no KV-cache and no generation loop — this module only defines the architecture's forward pass, not an inference/serving stack, which is intentionally out of scope for what was asked.
amalia/__init__.py
Re-exports AmaliaConfig and AmaliaForCausalLM (via __all__) so callers can do from amalia import AmaliaConfig, AmaliaForCausalLM instead of reaching into the submodules directly. This is the package's public API — everything else (RMSNorm, GroupedQueryAttention, etc.) is still importable, but isn't part of the "supported" surface.
main.py
A minimal runnable example: build a config, build the model with random weights, run one forward pass on random token ids, print the output shape. It exists purely to prove the wiring works end to end (uv run main.py) and to double as copy-paste-able usage code.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file amalia-0.5.1.tar.gz.
File metadata
- Download URL: amalia-0.5.1.tar.gz
- Upload date:
- Size: 27.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.18 {"installer":{"name":"uv","version":"0.9.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e1c40169a6431d67ef7eb955ba1010e1b5f31b6b2ecb51bb25df49e6b356ba51
|
|
| MD5 |
3eaee3444d5a72435013fd08f09cbdef
|
|
| BLAKE2b-256 |
6999d575181d706e54a01e61955a867965d91dacfb0aa3955459a7eb274311f0
|
File details
Details for the file amalia-0.5.1-py3-none-any.whl.
File metadata
- Download URL: amalia-0.5.1-py3-none-any.whl
- Upload date:
- Size: 7.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.18 {"installer":{"name":"uv","version":"0.9.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7a47650ac3c6af749d1acd043239a27b21a6af28321fcad44e01c012f992ba9c
|
|
| MD5 |
89c8052e1e1efcb3d76f7202a3e391ef
|
|
| BLAKE2b-256 |
934299c87119741b32fca6e2961f6f9f964901416b0a577c3a6de0571a662de0
|