edgenote
Prompt edge-pinning framework for LLMs. Mitigates the "Lost in the Middle" attention degradation problem by anchoring critical constraints and high-scoring context at prompt boundaries.
The Problem
Transformer architectures exhibit U-shaped attention distributions (Liu et al., 2023). Information positioned at the beginning (primacy) and end (recency) of the context window is retrieved reliably, while critical data in the middle suffers severe recall degradation.
edgenote structures prompts defensively:
- Edge Pinning: Anchors high-priority instructions, constraints, and facts to both the head and tail.
- U-Shaped Interleaving: Organises ranked documents so that top-scoring chunks occupy the high-attention edges while lower-ranked data remains in the middle.
- Exact Token Budgeting: Measures real token lengths and evicts or compresses low-priority middle context when limits are exceeded.
Empirical Benchmark
Multi-document needle-in-a-haystack retrieval evaluation across context positions:
| Position in Context | Baseline Prompt | edgenote |
Overhead |
|---|---|---|---|
| Start (0.0) | CORRECT | CORRECT | +44 tokens |
| 25% | WRONG | CORRECT | +44 tokens |
| Middle (0.5) | WRONG | CORRECT | +44 tokens |
| 75% | WRONG | CORRECT | +44 tokens |
| End (1.0) | CORRECT | CORRECT | +44 tokens |
Installation
pip install edgenote
Quick Start
from edgenote import Session
s = Session()
s.pin("Hardware budget is strictly capped at $185,000.", label="Constraint")
s.add("Proposal A: Liquid cooling loop overhaul ($240,000)", relevance=0.7)
s.add("Proposal B: High-density compute cluster ($180,000)", relevance=0.92)
result = s.render("Which proposal satisfies our constraints?")
print(result.text)
messages = result.to_messages()
Core Capabilities
1. Model-Accurate Token Counting
Pass any standard model identifier to bind the exact tokenizer backend:
from edgenote import Session
s = Session(token_counter="gpt-4o")
s = Session(token_counter="meta-llama/Meta-Llama-3-8B-Instruct")
Custom counting callables are also accepted:
import tiktoken
from edgenote import Session
enc = tiktoken.encoding_for_model("gpt-4o")
s = Session(token_counter=enc.encode)
2. Automated Re-ranking
Automatically score and reorder retrieved documents at render time without manual relevance labels:
from edgenote import Session
s = Session(reranker="cross-encoder")
s.add("Quarterly financial filings")
s.add("Personnel roster and team structure")
s.add("Enterprise procurement guidelines")
result = s.render("What were the fourth-quarter operating expenditures?")
Pure-Python zero-overhead rankers are also available:
from edgenote import Session
from edgenote.reranker import BM25Reranker, TFIDFReranker
s = Session(reranker=BM25Reranker())
s = Session(reranker=TFIDFReranker())
3. Context Compression
Summarize low-priority context when token limits are reached instead of dropping text:
from edgenote import Session
from groq import Groq
client = Groq()
def summarize(text: str, max_tokens: int) -> str:
response = client.chat.completions.create(
model="llama3-8b-8192",
messages=[
{"role": "system", "content": f"Summarize in {max_tokens} tokens or fewer."},
{"role": "user", "content": text},
],
max_tokens=max_tokens,
)
return response.choices[0].message.content
s = Session(compressor=summarize)
s.add("Long technical specification...")
result = s.render("Summarize system latency limits", budget=2048)
4. Text Chunking
Split large documents before ingestion:
from edgenote.chunker import StructuralChunker, SemanticChunker
chunker = StructuralChunker(max_tokens=512, overlap_tokens=64)
chunks = chunker.chunk(document_text)
semantic_chunker = SemanticChunker(threshold=0.5)
semantic_chunks = semantic_chunker.chunk(document_text)
5. Framework Integrations
LangChain
from edgenote.integrations import from_langchain
result = from_langchain(
docs,
query="What is the operating budget?",
pins=["Strictly cite sources using document headers."],
budget=4000,
)
messages = result.to_messages()
LlamaIndex
from edgenote.integrations import from_llamaindex
result = from_llamaindex(
nodes,
query="Synthesize quarterly performance metrics.",
pins=["Output format: Markdown table."],
)
Generic Dictionaries
from edgenote.integrations import from_dicts
docs = [
{"text": "Annual recurring revenue reached $12M", "relevance": 0.95, "source": "Finance"},
{"text": "Total headcount expanded to 120", "relevance": 0.40, "source": "HR"},
]
result = from_dicts(docs, query="Provide financial summary")
CLI & Model Cache
edgenote provides a command-line interface for verification and pre-caching neural weights:
# Verify installation and active backends
edgenote
# Run live prompt assembly demonstration
edgenote --demo
# Pre-cache weights for air-gapped environments
edgenote download bge-small
Author
- LinkedIn: Md Tareq Shah Alam
- Email: tareqshah.027@gmail.com
- License: Proprietary (All Rights Reserved)
Release files for edgenote 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| edgenote-0.1.3.tar.gz | 58.8 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| edgenote-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 58.8 MB
Release files / edgenote-0.1.3.tar.gz
| Download URL | edgenote-0.1.3.tar.gz |
|---|---|
| Size | 58.8 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ddda9d6b2fab3b33754c2d47b3490b4dcaa21498ef2c0ae2ba55c86901075418
|
|
BLAKE2b-256 checksum How to use checksums |
ba156e38338ace01e9dc2f33b9e5c0c349653ab6240255b87528fd38f96decff
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|
Release files / edgenote-0.1.3-py3-none-any.whl
| Download URL | edgenote-0.1.3-py3-none-any.whl |
|---|---|
| Size | 36.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
db0008050b0749014e93efeea9adf960146761d8444044259b1440025545cd60
|
|
BLAKE2b-256 checksum How to use checksums |
d7ad91107cdaf3f53acb21312f94e849164e83ef959542ffb2dde6a573f20444
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|