CNBE-32: Chinese Native Binary Encoding — Structured 32-bit CJK encoding for AI and hardware
Project description
CNBE-32
Chinese Native Binary Encoding
A 32-bit encoding that embeds the structural semantics of Chinese characters (radical, stroke count, and structure type) directly into binary, enabling CPUs and AI to natively understand Chinese.
A structured 32-bit encoding for 97,686 CJK characters that embeds radical, stroke count, and structure type directly into the encoding space.
[ Quick Start ] [ Key Experiments ] [ Tech Stack ] [ How to Contribute ] [ 中文 ] [ English ]
Architecture Panorama
graph TD
A[Chinese Character Input] -->|Radical/Stroke Count/Structure| B(CNBE-32 Encoder)
B -->|32-bit Binary| C{RISC-V Custom Instruction Set}
C -->|cnhe.map / cnhe.cmp| D[Hardware Layer: Spike/QEMU/FPGA]
D --> E[System Layer: Full Chinese Shell]
E --> F[Chinese BASIC / JEPA Semantic Engine]
Vision & Mission
Inspired by the Digital China 2035 strategy, CNBE-32's goal is:
To let every Chinese speaker seamlessly enter the AI era through their native language.
This is a mature system with a complete closed-loop, but as an early-stage exploration in the field of native Chinese computing, it remains in an open research phase. In the AI Agent era, the dreams of previous generations of scientists about full-Chinese computer systems finally have a chance to be realized.
Table of Contents
- Architecture Panorama
- Vision & Mission
- Code Quick Look
- Why CNBE?
- JEPA Exploration
- Cognitive Equity
- Key Experiments
- Key Insights I
- Experimental Limitations & Future Directions
- Tech Stack
- AI Agent Driven / AI Factory
- Quick Start
- Project Structure
- Roadmap
- How to Contribute
- Disclaimer
- License
Code Quick Look
Core Idea: Transform Chinese characters into 32-bit integers containing radical, stroke count, and structure type — letting the machine "see" the glyph directly.
CJK Character Mode (v6.0 Final)
Bit: 31 24 23 19 18 15 14 4 3 0
+----------------+--------+--------+------------------------+-------+
| Radical (8bit)|Stroke(5)|Struct(4)| Glyph Index (11bit) | Ext(4)|
+----------------+--------+--------+------------------------+-------+
| Field | Bit Range | Description | Range |
|---|---|---|---|
| Radical | [31:24] |
214 Kangxi radicals + 41 extensions | 0-255 |
| Stroke Count | [23:19] |
Number of strokes | 1-31 |
| Structure Type | [18:15] |
Structural composition type | 9 types (single/left-right/top-bottom/enclosure, etc.) |
| Glyph Index | [14:4] |
Intra-group index | 20,902 basic CJK characters |
| Extension | [3:0] |
Traditional/Simplified, ancient/modern, dialect, reserved flags | Reserved |
Encoding Examples
| Character | Unicode | CNBE-32 Encoding | Radical (ID) | Stroke Count | Structure Type |
|---|---|---|---|---|---|
| 一 (one) | U+4E00 | 0x01080000 |
一 (1) | 1 | Single (独体) |
| 汉 (Chinese) | U+6C49 | 0x0F288101 |
氵 (water, 15) | 5 | Left-Right (左右) |
| 国 (country) | U+56FD | 0x1F400B0B |
囗 (enclosure, 31) | 8 | Full Enclosure (全包围) |
| 明 (bright) | U+660E | 0x48400801 |
日 (sun, 72) | 8 | Left-Right (左右) |
Important: This is Not Base32
CNBE-32 is not a "Chinese-localized" or "character-replacement" version of Base32.
| Dimension | Base32 | CNBE-32 |
|---|---|---|
| Encoding target | Arbitrary binary data | 97,686 CJK characters themselves |
| Code space | Fixed 32 letters | Structured 32-bit bitfield (radical, stroke, structure) |
| Goal | Data compression / transmission | Let machines "understand" character semantics |
| Target audience | Human-readable (transcription) | AI models, CPU instruction sets, OS kernels |
In one sentence: Base32 turns data "into letters", CNBE-32 turns characters "into semantics".
Who is it for?
- AI models: Structured prior knowledge input (radical=spatial anchor, stroke=discrete feature, structure=spatial relationship)
- CPU instruction sets:
cnhe.map/cnhe.extract/cnhe.cmpoperate at the hardware level - OS kernels: Filenames, paths, system messages natively support CNBE-32 encoding
- Not recommended: URL transmission, database primary keys, human transcription (use Base64/Base32)
Why CNBE?
| Dimension | Unicode / UTF-8 | CNBE-32 |
|---|---|---|
| Objective | Character display and exchange | AI understanding and hardware acceleration |
| Encoding Method | Lookup table (Flat ID) | Semantic structuring |
| Machine Cognition | Identifies the character | Understands structural composition |
| AI Compatibility | Learns from data | Provides structural priors |
10 cross-domain validations passed (incl. LLM LoRA training): Linguistics, Ecology, Meteorology, Finance, Biology, Physics, Sociology, Pre-training, Mathematics
JEPA Exploration
CNBE is not a patch for today's Transformers, but foundational infrastructure for tomorrow's JEPA.
Yann LeCun's JEPA emphasizes prediction in representation space — and CNBE provides exactly the most structured representation space:
- Radical = Spatial Anchor: Characters sharing the same radical naturally cluster in binary space
- Stroke Count = Discrete Feature: Provides fine-grained morphological differentiation
- Structure = Spatial Relationship: Left-right, top-bottom, enclosure, etc. directly map to topological relationships
Completed JEPA validations: v9 tree structure prediction + v10 cross-9-domain generalization
Cognitive Equity
The underlying logic of modern computers (from instruction sets to OS kernels) is built entirely on English/Latin alphabets. This creates a cognitive barrier for non-native English speakers who must first translate their thoughts before performing low-level development.
The ultimate significance of CNBE-32 is to enable Chinese speakers to define underlying logic directly through their native linguistic thinking, breaking down professional vocabulary barriers and achieving true technological cognitive equity.
In the AI era, every Chinese user — regardless of age, education level, or professional background — should be able to engage in deep dialogue with AI, define rules, and even write underlying logic using their native language.
Core Performance Overview
| Metric | Value | Equity Value Explanation |
|---|---|---|
| Small model (<1B) comprehension improvement | +81% (48%→87%) | Edge devices can achieve high-quality Chinese comprehension without cloud connectivity, breaking the compute monopoly of large tech companies |
| Medium model (1-7B) improvement | +9% ~ +17% | Mid-range mobile chips can smoothly run complex Chinese tasks without relying on high-end GPUs |
| Large model (>7B) benefit | ~0% (diminishing returns) | Validates that large models don't need this encoding; resources should be prioritized for small-to-medium intelligence scenarios |
| Hardware lookup extreme latency | 0.8 ns (x86) / 1 Cycle (FPGA) | Ultra-fast response for real-time interaction; suitable for low-frequency, low-power embedded chips |
| Minimal memory footprint | Only 81.6 KB (SRAM/BRAM) | Fits easily into any L1/L2 cache or on-chip storage without external DRAM, reducing BOM cost |
| Encoding semantic density | 32 bits containing radical/stroke/structure | Single encoding equivalent to dozens of text annotation tokens, greatly reducing learning and inference overhead for small models |
| CJK coverage breadth | 97,686 characters | Covers ancient texts, rare names, and dialect characters, ensuring cultural diversity isn't marginalized in the AI era |
| Hard-task rare character handling | +17.4 pp (vs Unicode) | Dominates traditional encoding in traditional/variant/chemical equation scenarios, ensuring professional knowledge equity |
| Lookup collision rate | 0% (full coverage verified) | Zero-ambiguity lookup, ensuring stability and reliability on edge devices |
Key Experiments
Small Model, Big Improvement (v2)
Hypothesis: Structured encoding compensates for insufficient small model parameters. Method: Qwen 3.5 0.8B, CNBE vs standard input.
| Input | Accuracy | Improvement |
|---|---|---|
| Standard input | 48% | -- |
| CNBE-32 | 87% | +81% |
CNBE Surpasses Unicode (v6.5.2)
Hypothesis: Structured bit fields carry more semantic information than Unicode code points. Method: Gemma 4B Chinese hard tasks.
| Input | Accuracy |
|---|---|
| Unicode | 26.1% |
| CNBE-32 | 43.5% |
Conclusion: A brand-new encoding without prior training outperforms the 30-year standard on first attempt (+17.4 pp).
Full Chinese Operating System (v8.4)
- Full Chinese Shell (output/get encoding/compare commands)
- Chinese BASIC interpreter (7 keywords)
- Text editor (built-in Tao Te Ching, 205 lines)
- RISC-V custom instructions:
cnhe.map/cnhe.extract/cnhe.cmp
Mathematical Reasoning Foundation (v10.8)
Method: TinyGPT on odd/even/prime/sequence reasoning tasks comparing 4 encodings.
| Task | CNBE Loss | OneHot Loss | Winner |
|---|---|---|---|
| Odd/Even | 0.3174 | 0.3427 | CNBE |
| Prime | 0.3894 | 0.5061 | CNBE |
| Sequence | 1.0726 | 1.2344 | CNBE |
Complete Experimental Data (v1~v10)
Click to expand v1~v10 core experiment overview
| Version | Validation Dimension | Model / Platform | Core Metric | Key Conclusion |
|---|---|---|---|---|
| v1 | Zero-shot single character understanding | Qwen 0.8B | 200 characters, 100% effective | Encoding is inherently semantically interpretable |
| v2 | Small model sentence understanding | Qwen 0.8B | 48% → 87% (+81%) | Structured encoding provides significant compensation for small models |
| v3 | Annotation format optimization | Qwen 0.8B | Character-by-character full annotation 87% effective | Optimal format: character-by-character full annotation |
| v4 | Long text (paper-level) | Qwen 0.8B | 90.9% → 100% | Effective in long-text scenarios, eliminates ambiguity |
| v5 | Multi-model horizontal comparison | 7 models | <1B: +81%; 1-7B: +9~17%; >7B: ~0% | Diminishing marginal returns |
| v6 | Unicode hard task comparison | Gemma 4B | Unicode 26.1% vs CNBE 43.5% | CNBE > Unicode (+17.4 pp) |
| v7 | RISC-V hardware implementation | C / QEMU / Spike / FPGA | x86 0.8 ns → FPGA 1 Cycle | Complete hardware path closed-loop |
| v8 | Full Chinese operating system | RISC-V QEMU | Chinese Shell + BASIC + Tao Te Ching editor | Encoding can seamlessly integrate into OS underlying layer |
| v9 | JEPA tree structure prediction | JEPA architecture | Error 0.0899 → 0.000001 | Extremely strong high-noise temporal feature extraction |
| v10 | Cross-9-domain generalization | Multi-domain | Mathematics wins; typhoon error −19% | Effective across mathematics/physics/biology/finance and other domains |
Click to expand v1~v10 detailed experimental data
| Version | Sub-item / Task | Test Environment | Specific Data Metrics | Conclusion / Notes |
|---|---|---|---|---|
| v1 | Single character radical/stroke/structure extraction | Qwen 0.8B | 200 Chinese characters, 100% zero-shot effective | Proves encoding space IS semantic space |
| v2 | Chinese sentence understanding | Qwen 0.8B | Text input 48% → CNBE 87% | Accuracy improvement 39 pp |
| v3 | Encoding format ablation experiment | Qwen 0.8B | Character-by-character 87% > segmented 60% > compact 50% | Optimal: 中(丨,4 strokes, single) |
| v4 | Paper-level semantic understanding | Qwen 0.8B | 90.9% → 100% | Complements small model long-context reasoning shortcomings |
| v5a-5.9 | 7-model horizontal comparison | 0.8B~20B | Domestic 2B 90%; 8B+ approaches 0 | Less compute power = more important structural priors |
| v6.3-6.5 | Numerical format optimization | Qwen 0.8B | Format F (bare numbers) optimal | Hardware recommends bare number input |
| v6.5.2 | CNBE vs Unicode | Gemma 4B | Unicode 26.1% vs CNBE 43.5% | Outperforms thirty-year industry standard on first attempt |
| v7.0 | C language benchmark | x86-64 | Single lookup 0.8 ns | Software performance baseline established |
| v7.0.1 | RISC-V cross-compilation | QEMU | Single lookup 2.5 ns | Validates RISC-V portability |
| v7.1.1 | Instruction integration | Spike | map(2 cycles) / extract(1) / cmp(3) |
Three Custom-0 instruction behaviors verified |
| v7.2 | FPGA logic synthesis | Verilog+BRAM | Single cycle lookup complete | 81.6 KB table entries fit BRAM resources |
| v8.4 | Full Chinese system | RISC-V QEMU | Shell commands + BASIC 7 keywords + Tao Te Ching | "Full Chinese computing" feasibility validated |
| v9.0 | Tree growth JEPA | JEPA | CNBE 86% better than Raw | Structured encoding improves abstract representation |
| v9.1 | Typhoon lifecycle | JEPA | 0.089981 → 0.000001 | Error reduced by 4 orders of magnitude |
| v10.3 | Typhoon Bavi path | Meteorological model | 216 km → 174 km | Actual path prediction accuracy improved 19% |
| v10.4 | Protein Q3 structure | Bioinformatics | OH 44.6% vs CNBE 41.0% | Slightly below OH; biological sequence still has optimization room |
| v10.5 | Black hole gravitational field | Physics simulation | R² 0.60-0.77 | Good performance in physics field simulation |
| v10.7 | TinyGPT frozen embedding | TinyGPT | Learned 1.3653 vs CNBE 1.4568 | Frozen embedding performance close to learned embedding |
| v10.8 | Mathematical reasoning foundation | TinyGPT | Odd/Even(0.3174<0.3427) Prime(0.3894<0.5061) Sequence(1.07<1.23) | Universally better than One-Hot |
| LLM | CNBE knowledge LoRA fine-tuning | Qwen3.5-0.8B | 5000 steps, loss 0.6424, 500-steps 0.7524 | Knowledge injection feasible, edge deployment validated |
Complete Evidence Chain Logic Closure
| Stage | Corresponding Version | Logical Role |
|---|---|---|
| Semantic Validity | v1 ~ v4 | Prove encoding itself contains semantics |
| Comparative Superiority | v5 ~ v6 | Prove encoding outperforms Unicode |
| Hardware Implementability | v7 | Prove from software to FPGA is feasible |
| System-level Compatibility | v8 | Prove encoding can support complete OS ecosystem |
| Cross-domain Generalization | v9 ~ v10 | Prove equally effective in physics/biology/finance and other domains |
Complete experimental data → docs/EXPERIMENTS.md
Click to expand v1-v10.8 complete experiment data
Table 1: CNBE-32 Core Experiment Overview (v1~v10)
| Version | Dimension | Model / Platform | Key Metric | Key Conclusion |
|---|---|---|---|---|
| v1 | Zero-shot char understanding | Qwen 0.8B | 200 chars, 100% effective | Encoding inherently semantically interpretable |
| v2 | Small model sentence understanding | Qwen 0.8B | 48% → 87% (+81%) | Structured encoding compensates small models significantly |
| v3 | Annotation format optimization | Qwen 0.8B | Full char annotation 87% effective | Optimal format: per-character annotation |
| v4 | Long text (paper-level) | Qwen 0.8B | 90.9% → 100% | Effective in long-context scenarios |
| v5 | Multi-model comparison | 7 models | <1B: +81%; 1-7B: +9~17%; >7B: ~0% | Diminishing returns law |
| v6 | Unicode hard task comparison | Gemma 4B | Unicode 26.1% vs CNBE 43.5% | CNBE > Unicode (+17.4pp) |
| v7 | RISC-V hardware implementation | C/QEMU/Spike/FPGA | x86 0.8ns → FPGA 1 Cycle | Complete hardware path closed-loop |
| v8 | Full Chinese OS | RISC-V QEMU | Chinese Shell + BASIC + Dao De Jing | Encoding integrates seamlessly into OS |
| v9 | JEPA tree structure prediction | JEPA architecture | Error 0.0899 → 0.000001 | Powerful feature extraction for noisy data |
| v10 | Cross 9-domain generalization | Multi-domain models | Math wins, typhoon error -19% | Effective in math/physics/biology/finance |
Table 2: Detailed Experiment Data (v1~v10)
| Version | Sub-task | Environment | Specific Metric | Conclusion |
|---|---|---|---|---|
| v1 | Char radical/stroke/structure extraction | Qwen 0.8B | 200 chars, 100% zero-shot | Encoding space = semantic space |
| v2 | Chinese sentence understanding | Qwen 0.8B | Text 48% → CNBE 87% | +39pp absolute improvement |
| v3 | Encoding format ablation | Qwen 0.8B | Per-char 87% > segment 60% > compact 50% | Optimal: full per-char annotation |
| v4 | Paper-level semantic understanding | Qwen 0.8B | 90.9% → 100% | Fills small model long-context gap |
| v6.5.2 | CNBE vs Unicode | Gemma 4B | Unicode 26.1% vs CNBE 43.5% | Surpassed 30-year standard first try |
| v7.1.1 | Custom instruction integration | Spike | map(2 cycles)/extract(1)/cmp(3) | Three Custom-0 instructions verified |
| v7.2 | FPGA logic synthesis | Verilog+BRAM | Single cycle lookup | 81.6KB table fits BRAM |
| v8.4 | Full Chinese system | RISC-V QEMU | Shell + BASIC 7 keywords + Dao De Jing | Chinese computing feasibility proven |
| v9.0 | Tree growth JEPA | JEPA | CNBE 86% better than Raw | Structured encoding boosts abstraction |
| v9.1 | Typhoon lifecycle JEPA | JEPA | 0.089981 → 0.000001 | Error reduced 4 orders of magnitude |
| v10.3 | Typhoon Barijat path | Meteorological model | 216 km → 174 km | Path prediction accuracy +19% |
| v10.4 | Protein Q3 structure | Bioinformatics | OH 44.6% vs CNBE 41.0% | Slightly below OH, room for improvement |
| v10.5 | Black hole gravity simulation | Physics simulation | R² 0.60-0.77 | Physical field simulation performs well |
| v10.7 | TinyGPT frozen embedding | TinyGPT | Learned 1.3653 vs CNBE 1.4568 | Close to learned as frozen embedding |
| v10.8 | Math reasoning base | TinyGPT | Parity/Prime/Seq CNBE wins all | Comprehensive win over One-Hot |
Table 3: Evidence Chain Logic Closure
| Phase | Version | Role |
|---|---|---|
| Semantic validity | v1~v4 | Proves encoding contains semantics |
| Comparative superiority | v5~v6 | Proves encoding > Unicode |
| Hardware feasibility | v7 | Proves software-to-FPGA path |
| System-level compatibility | v8 | Proves encoding supports full OS ecosystem |
| Cross-domain generalization | v9~v10 | Proves effective across multiple domains |
Full experiment data → docs/EXPERIMENTS.md
Key Insights: Large Models vs Small Models
Why do 8B+ large models show diminishing returns (~0%) from CNBE, while 0.8B small models achieve massive +81% improvement?
- Large Model Brute Force Aesthetics: Massive parameters can implicitly memorize Unicode through brute-force training, masking the structural flaws of the encoding
- Small Model Structural Priors: On compute-constrained edge devices, CNBE transforms glyph structure directly into computational priors
This is the breakthrough path for edge-side AI processing of Chinese.
Key Insights III: CNBE Encoding Knowledge LoRA Fine-Tuning
— Injecting CNBE-32 encoding knowledge into Qwen3.5-0.8B via LoRA
- LoRA knowledge injection works: 500 steps (22 min) + 5000 steps (4.14 h) with 25K diverse Chat Template data, loss from 0.7524 → 0.6424 (↓14.6%), augmentation artifacts eliminated
- Model understands encoding concepts: After fine-tuning, the model recognizes character radicals, stroke counts, and structure types, outputting CNBE-32 encoded information
- Minimal GPU requirements: RTX 4060 Ti (8GB) handles the entire pipeline, with peak memory usage of only 1.5GB
- Edge deployment validated: For the first time, CNBE-32 advances from inference-level semantic validation to training-level knowledge injection
- Complete cross-domain chain: From linguistics to finance to physics to biology to LLM training, CNBE's structured encoding is validated across encoding, hardware, OS, cross-domain prediction, and model fine-tuning
Full methodology → cnbe-llm training(demo)/
Experimental Limitations & Future Directions
We have faithfully documented failures and limitations in all experiments. The following are known boundaries disclosed directly in this README.
Known Limitations
| Experiment | Limitation | Future Direction |
|---|---|---|
| v5/v6 (LLM validation) | Some models (DeepSeek 8B / GPT-OSS 20B) showed empty responses or insufficient Chinese capability | Focus on Chinese-friendly small models like Qwen/Gemma |
| v6.5.3 (Hard task 0.8B) | Overall only 12.5%, CNBE and Unicode showed no difference | 0.8B model capability boundary; requires larger model validation |
| v9.0 (Tree growth) | Simulated environment, not real climate/economic data | Validate on real temporal data |
| v10.0/v10.1 (Financial backtesting) | A-share high-frequency trading costs (0.14%/trade) consumed all strategy returns; break-even point not reached | Pivot to low-frequency strategies (daily/weekly) to unlock predictive value |
| v10.4 (Protein) | Used simplified single-residue method, not standard sliding window; first contact with 30-year domain standard gap of 3.6 pp | Sliding window + CB513 dataset complete experiment |
| v10.5 (Black hole) | Single-variable input scenario (only r/Rs), continuous value KNN naturally precise; CNBE quantization introduces error | Multi-dimensional input scenario (with observation noise) validation |
| v10.6 (Sociology) | CNBE inferior to One-hot in strong classification feature scenarios (MSE 0.0124 vs OneHot 0.0019) | Field weighting, hierarchical encoding optimization |
| v10.7 (Pre-training) | Task too simple (13 token vocabulary), difference not statistically significant | Large-scale corpus, larger model validation |
Applicability Boundaries (Based on All Experimental Data)
| Scenario Type | CNBE Performance | Typical Domains | Reason |
|---|---|---|---|
| Multi-dimensional continuous value + structured temporal | ✅ Significantly better than baseline | Meteorology, ecology, finance, mathematics | Bit-field structured encoding naturally matches |
| Strong classification features | ❌ Inferior to One-hot | Sociology (8 regions + 4 time periods) | Bit-field mixed encoding cannot distinguish classification field weights |
| Single-variable deterministic systems | ⚠️ Equal to Raw | Physics (gravitational field) | Continuous value single-variable scenario Raw is optimal |
| Zero-shot unfamiliar domains | ⚠️ Close to domain standard | Biology (protein) | First attempt approaches 30-year optimized standard |
| Pattern recognition tasks | ✅ Universally better than One-hot | Mathematical reasoning | Structured encoding matches pattern recognition |
Tech Stack
Application Layer: Chinese BASIC interpreter + Text Editor + Tao Te Ching
System Layer: Full Chinese Shell + CNBE Runtime (map/extract/cmp)
Hardware Layer: RISC-V 1GHz + 1GB RAM (QEMU + Spike)
Instruction Layer: cnhe.map / cnhe.extract / cnhe.cmp
Encoding Layer: 32-bit CJK Structured Bit Fields (Radical/Stroke/Structure)
AI Agent Driven / AI Factory
This is a project that was previously impossible to complete, but is destined to be born in the AI era.
| Past | Present |
|---|---|
| 97,686 Chinese character annotations required thousands of linguist man-years | AI Agent assisted automated annotation |
| Full-stack validation required top-tier teams for years | LLM-assisted code generation + validation |
| Single-team siloed development | Open source community collaborative exploration |
The dreams of scientists from the last century finally have a chance to be realized in the AI Agent era.
Quick Start
Environment Requirements
- Python 3.8+
- numpy, torch, scikit-learn (for experiment reproduction)
Install Python SDK
pip install numpy torch scikit-learn
Usage Example
import sys; sys.path.insert(0, 'src')
from cnbe32 import encode_cnbe, hamming_distance
code_ming = encode_cnbe(72, 8, 1) # 明 (bright) = 日(sun, 72) + 8 strokes + left-right structure
code_an = encode_cnbe(72, 9, 1) # 暗 (dark) = 日(sun, 72) + 9 strokes + left-right structure
print(hamming_distance(code_ming, code_an))
Run RISC-V Simulator
cd hardware/simulator
gcc -o cnhe_sim cnhe_sim.c -Wall -O2 && ./cnhe_sim
Launch Full Chinese Operating System (QEMU)
# Ubuntu dependencies
sudo apt-get install -y gcc-riscv64-linux-gnu qemu-system-misc
cd v84_riscv_os_full
make all && make run
Reproduce Experiments
cd v10_8_math_reasoning && python run_v108.py
cd v10_3_typhoon && python v10_3_typhoon.py
Application Scenarios
CNBE-32 is designed for AI-era Chinese computing infrastructure, not as a general-purpose encoding tool.
| Scenario | Suitability | Description |
|---|---|---|
| AI model structured input | Recommended | Provides radical/stroke/structure priors, improves small model comprehension |
| RISC-V hardware acceleration | Recommended | Custom instructions operate directly on encoding bitfields |
| Chinese-native OS | Recommended | Native support for filenames, paths, system messages |
| Chinese compiler/BASIC | Recommended | Direct encoding operations at the language level |
| Data compression/obfuscation | Not recommended | Semantic encoding, not compression |
| URL transmission | Not recommended | Characters broken by %-encoding in URLs |
| Human transcription | Not recommended | Visually similar characters cause errors |
| Database primary keys | Use with caution | 32-bit integers are storable but CNBE is not a unique identifier |
Project Structure
CNBE-32-Chinese-Native-Binary-Encoding/
|-- docs/specification/ # Encoding specification
|-- docs/EXPERIMENTS.md # Experiment overview
|-- docs/VISION.md # Strategic vision
|-- src/cnbe32/ # Python SDK
|-- include/cnbe32.h # C header file
|-- data/ # Encoding database
|-- tests/ # Test suite
|-- tools/ # Development tools
|-- bindings/rust/ # Rust bindings
|-- hardware/ # RISC-V simulator
|-- v9_jepa_tree/ # JEPA experiments (v9)
|-- v10_5~v10_8/ # Cross-domain experiments (v10)
|-- v84_riscv_os_full/ # Chinese OS prototype
|-- results/ # White papers (41 documents)
|-- LICENSE # Mulan License
Roadmap
| Phase | Status | Content |
|---|---|---|
| Encoding & semantic validation | Completed | v1-v6 CJK encoding design |
| Hardware & system | Completed | v7-v8 RISC-V + Chinese OS |
| Complex prediction validation | Completed | v9-v10 9-domain validation |
| AI compiler | Planned | Chinese natural language → machine code |
| Edge AI integration | Planned | Edge AI default standard |
| Ecosystem collaboration | Vision | Open source community + industry standards |
How to Contribute
Current Directions Most Needing Community Support
- Chinese BASIC interpreter optimization - improve lexical analyzer
- RISC-V lookup logic acceleration - optimize 81.6 KB L2 Cache hit rate
- JEPA architecture extension experiments - more physics/biology system tests
- Frontend visualization tools - Web interface showing encoding decomposition process
| Level | Direction |
|---|---|
| Low barrier | Encoding dictionary / Test cases / Documentation |
| High barrier | RISC-V pipeline / FPGA / LLM adaptation / Compiler |
See CONTRIBUTING.md for details
Disclaimer
v10.x stage financial time series (US stocks / A-shares) backtesting is solely for validating CNBE-32's feature extraction and structured prior capabilities in high-noise, non-stationary time series data, and does not constitute any investment advice.
License
Mulan Permissive Software License v2 (Mulan PSL v2)
Let Chinese speakers enter the AI era through their native language.
From the "Digital China 2035" vision to AI Agent era engineering practice.
Born for Chinese AI ecosystem — from encoding to hardware, from single character to operating system.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file cnbe32-1.0.0.tar.gz.
File metadata
- Download URL: cnbe32-1.0.0.tar.gz
- Upload date:
- Size: 42.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c28491cda973a88833bdc8a8c87335a363133daac0488e81223b0a39958e8ebd
|
|
| MD5 |
3122b9df462ec4153703e605fb7449ce
|
|
| BLAKE2b-256 |
bf623fc102856adf2009364b749b3509c42ffc19c8f29b9864bb490a0d6a1527
|
File details
Details for the file cnbe32-1.0.0-py3-none-any.whl.
File metadata
- Download URL: cnbe32-1.0.0-py3-none-any.whl
- Upload date:
- Size: 22.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d3f91912e7dd7228926f8de25cf631a7f818ff825782aaa996f82a644c4b7787
|
|
| MD5 |
5a250dc593dcb2453b676804afb13214
|
|
| BLAKE2b-256 |
0d8139f8fb5a29de5e1b29b06429d3b7e2d18815a549dd998367d3b841317b5a
|