Skip to main content

qwen38-flash-next

在本地一键运行 Qwen3.8-Flash-Next(125B MoE,6B 激活)UD-IQ1_S 量化版。

执行后自动完成:查找或安装 llama.cpp 运行时(qwen4exp 架构需要未合入的 PR #27742 构建)→ 查找或下载模型(默认走 hf-mirror.com,国内网络友好)→ 启动本地服务 → 进入交互式对话。已存在的模型和运行时会被直接复用,不重复下载/编译。

安装

pip install qwen38-flash-next        # PyPI
# 或本地源码
uv tool install ./qwen38-flash-next  # 推荐,隔离环境

使用

qwen38-flash                 # 交互式对话
qwen38-flash --say "你好"     # 单发一句(脚本/测试用)
qwen38-flash --download-only # 只下载模型

首次运行需要约 74GB 磁盘与下载时间;48GB 内存机器默认 -ngl 0 纯 CPU 运行(Metal 会因超大嵌入层 OOM)。

交互命令:/exit 退出,/clear 清空上下文,/save 文件 保存对话。

常用选项

参数 / 环境变量 说明
--model PATH / QWEN38_MODEL 指向模型首分片 ...00001-of-00003.gguf,跳过查找与下载
--server-bin PATH / QWEN38_LLAMA_SERVER 指定支持 qwen4exp 的 llama-server 二进制
--port 8080 / --ctx 8192 / --ngl 0 服务端口 / 上下文 / GPU 层数
HF_ENDPOINT 模型下载源,默认 https://hf-mirror.com

模型与运行时路径在首次就绪后写入 ~/.qwen38-flash/config.json,后续启动直接复用。

查找顺序(逐级复用)

模型:--model/环境变量 → 配置文件 → 常见位置(HF 缓存、ModelScope 缓存、~/models、本仓库工作区)→ 自动下载。 运行时:--server-bin/环境变量 → 配置文件 → PATH 中的 llama-server(会校验能否加载模型)→ 自动克隆 PR #27742 源码编译(需 git + cmake + Xcode CLT)。

Release files for qwen38-flash-next 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for qwen38-flash-next 0.1.1
File Size Uploaded
qwen38_flash_next-0.1.1.tar.gz 9.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for qwen38-flash-next 0.1.1
File Interpreter ABI Platform
qwen38_flash_next-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 20.1 kB

Release files / qwen38_flash_next-0.1.1.tar.gz

Download URL qwen38_flash_next-0.1.1.tar.gz
Size 9.4 kB
Tags Source
SHA-256 checksum
How to use checksums
7b72255c5f29bf2d11843a5002970d99ebf1f079c22d080764db3054e430e9d3
BLAKE2b-256 checksum
How to use checksums
779a16a4ca6b16b28755b50d29715cee53af40b76d315134f481792c32a59ddb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / qwen38_flash_next-0.1.1-py3-none-any.whl

Download URL qwen38_flash_next-0.1.1-py3-none-any.whl
Size 10.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1b6ae3167dffb3bb1d3824eb7d4947a9d42029b689cc5cea309770c903a09ac7
BLAKE2b-256 checksum
How to use checksums
208fd153e0b84f04191d6cce7ea4d12f55adf8f19c74d9876ba4b340f932be2f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page