fitzyracing-transformers-stream-generator
Part of 360 Bench, tested fixes for abandoned PyPI packages.
This is a fork of transformers-stream-generator by LowinLi, published as a drop-in replacement that works with current transformers. The upstream package (~360k downloads a month, required by the remote code of Qwen(-1) models, among others) has not been released since 0.0.5 (March 2024), and on transformers 4.57 and newer it can't even be imported. The import name is still
transformers_stream_generator, so no code changes are needed. All credit for the original library goes to its author; it remains available under the same MIT license.If a fixed
transformers-stream-generatorrelease appears on PyPI, prefer it and switch back.
What's fixed in this fork
Based on upstream transformers-stream-generator==0.0.5 (tag
0.0.5). The API is unchanged.
0.0.5 copies the generate() method of transformers 4.26 and builds on transformers internals
that have since changed. What that meant, tested with a tiny GPT-2:
| transformers | 0.0.5 | this fork |
|---|---|---|
| 4.26 to 4.40 | works | unchanged: runs exactly the same 0.0.5 code |
| 4.41 to 4.56 | imports, but after init_stream_support() every model.generate() call fails (TypeError/AttributeError), streaming or not |
works |
| 4.57 and 5.x | ImportError: cannot import name 'BeamSearchScorer' (4.57) / 'DisjunctiveConstraint' (5.x) on import |
works |
On transformers 4.41 and newer the fork:
- imports without the classes transformers removed (they were only used by the old code);
- streams
model.generate(..., do_stream=True)through transformers' own publicgenerate(streamer=...), running generation in a background thread and yielding each step's token ids (a tensor of shape(batch_size,)on the input's device), exactly like 0.0.5. With the sameseedit yields the same tokens as non-streaming sampling. Qwen(-1)'schat_stream(), which callsNewGenerationMixin.generatewith aStreamGenerationConfig(do_stream=True), works the same way. If you stop reading early (break,generator.close()), generation stops too; errors are raised in your loop; - passes every other
generate()call (greedy, sampling, beam search, ...) to transformers' owngenerate(), soinit_stream_support()no longer breaks them; - keeps 0.0.5's conventions: stream mode always samples (as 0.0.5 did, even with
do_sample=False; see #1), andgenerate()seeds the random generators withseed=0unless you pass anotherseed(seed=-1leaves them alone).
What doesn't work on transformers 4.41+, with a clear error instead of a crash:
- Beam search can't stream.
do_stream=Truewithnum_beams > 1raisesValueError: do_stream=True only supports sampling with num_beams=1 ...(0.0.5 silently returnedNone, see #3). Beam search withoutdo_streamworks. model.sample_stream()(the internal sampler, which 0.0.5 also exposed) relies on transformers internals that no longer exist, so calling it raisesNotImplementedErrortelling you to usegenerate(..., do_stream=True). Assigning it, as Qwen(-1)'s remote code does, still works.
As a side effect, the new stream also handles a list of several eos_token_ids correctly
(#10) and batches of more
than one prompt. On transformers < 4.41 these 0.0.5 limitations are kept as they were, because
that code is deliberately left untouched.
Upstream issues: #15 (no fix
PR existed). Note that Qwen(-1)'s own remote code may have other incompatibilities with recent
transformers; this fork only makes transformers_stream_generator itself work.
Packaging: pyproject.toml replaces setup.py; dependencies are unchanged
(transformers>=4.26.1). The license metadata now says MIT only, matching the LICENSE file
(0.0.5's classifiers also listed Apache).
Install
pip install fitzyracing-transformers-stream-generator
Switching from transformers-stream-generator
This distribution installs the same transformers_stream_generator package as the original, so
your code stays the same. Uninstall the original first, then install the fork:
pip uninstall -y transformers-stream-generator
pip install fitzyracing-transformers-stream-generator
The order matters. Both distributions own the same files, so if you install the fork first and
uninstall the original afterwards, pip deletes the shared files and the import breaks. If that
happens, run pip install --force-reinstall --no-deps fitzyracing-transformers-stream-generator.
In requirements.txt / pyproject.toml, replace transformers-stream-generator (or
transformers_stream_generator) with fitzyracing-transformers-stream-generator.
If you get it through another package
pip cannot replace a dependency with a differently named package. If a dependency requires
transformers-stream-generator, install the fork alongside it and then remove the original's
files:
pip install fitzyracing-transformers-stream-generator
pip uninstall -y transformers-stream-generator
pip install --force-reinstall --no-deps fitzyracing-transformers-stream-generator # restore the files the uninstall removed
Afterwards pip check reports <package> requires transformers-stream-generator, which is not installed; that is expected. Repeat the steps if a later install pulls the original back in.
uv users can do this properly with an override that drops the original:
# pyproject.toml
[project]
dependencies = ["fitzyracing-transformers-stream-generator", "...the package that depends on it..."]
[tool.uv]
override-dependencies = ["transformers-stream-generator; sys_platform == 'never'"]
(or uv pip install --override overrides.txt ... with that same line in overrides.txt).
Tests
pip install -e ".[test]"
pytest
tests/test_fork_fixes.py uses hf-internal-testing/tiny-random-gpt2 (about 2 MB). It passes
on transformers 4.26, 4.32, 4.36, 4.38, 4.40, 4.41, 4.45, 4.52, 4.57 and 5.19 with this fork;
with 0.0.5 it fails on everything from 4.41 on.
transformers-stream-generator
Description
This is a text generation method which returns a generator, streaming out each token in real-time during inference, based on Huggingface/Transformers.
Web Demo
- original
- stream
Installation
pip install transformers-stream-generator
Usage
- just add two lines of code before your original code
from transformers_stream_generator import init_stream_support
init_stream_support()
- add
do_stream=Trueinmodel.generatefunction and keepdo_sample=True, then you can get a generator
generator = model.generate(input_ids, do_stream=True, do_sample=True)
for token in generator:
word = tokenizer.decode(token)
print(word)
Example
Metadata
Release files for fitzyracing-transformers-stream-generator 0.0.6
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| fitzyracing_transformers_stream_generator-0.0.6.tar.gz | 19.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| fitzyracing_transformers_stream_generator-0.0.6-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 36.8 kB
Release files / fitzyracing_transformers_stream_generator-0.0.6.tar.gz
| Download URL | fitzyracing_transformers_stream_generator-0.0.6.tar.gz |
|---|---|
| Size | 19.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
6b89da6bc94c471bb96ad5157f8fde495ff42adc9e8508880731eae8b053a362
|
|
BLAKE2b-256 checksum How to use checksums |
8fb1424d61c0a21132009161ac368358ec024b48536f9089ccb568b14d8557a3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.5
|
Release files / fitzyracing_transformers_stream_generator-0.0.6-py3-none-any.whl
| Download URL | fitzyracing_transformers_stream_generator-0.0.6-py3-none-any.whl |
|---|---|
| Size | 17.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1f9a15dc0248ca11bd768c84ee4c20e8346eee5a1a06d25efc46b0c712e43179
|
|
BLAKE2b-256 checksum How to use checksums |
1e12dfb0e871ac706a7a3ad2e7d7703b5c1b7d6467c85992662d09caa38abcf3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.5
|