Skip to main content

edgenote

Prompt edge-pinning framework for LLMs. Mitigates the "Lost in the Middle" attention degradation problem by anchoring critical constraints and high-scoring context at prompt boundaries.


The Problem

Transformer architectures exhibit U-shaped attention distributions (Liu et al., 2023). Information positioned at the beginning (primacy) and end (recency) of the context window is retrieved reliably, while critical data in the middle suffers severe recall degradation.

edgenote structures prompts defensively:

  1. Edge Pinning: Anchors high-priority instructions, constraints, and facts to both the head and tail.
  2. U-Shaped Interleaving: Organises ranked documents so that top-scoring chunks occupy the high-attention edges while lower-ranked data remains in the middle.
  3. Exact Token Budgeting: Measures real token lengths and evicts or compresses low-priority middle context when limits are exceeded.

Empirical Benchmark

Multi-document needle-in-a-haystack retrieval evaluation across context positions:

Position in Context Baseline Prompt edgenote Overhead
Start (0.0) CORRECT CORRECT +44 tokens
25% WRONG CORRECT +44 tokens
Middle (0.5) WRONG CORRECT +44 tokens
75% WRONG CORRECT +44 tokens
End (1.0) CORRECT CORRECT +44 tokens

Installation

pip install edgenote

Quick Start

from edgenote import Session

s = Session()
s.pin("Hardware budget is strictly capped at $185,000.", label="Constraint")
s.add("Proposal A: Liquid cooling loop overhaul ($240,000)", relevance=0.7)
s.add("Proposal B: High-density compute cluster ($180,000)", relevance=0.92)

result = s.render("Which proposal satisfies our constraints?")
print(result.text)

messages = result.to_messages()

Core Capabilities

1. Model-Accurate Token Counting

Pass any standard model identifier to bind the exact tokenizer backend:

from edgenote import Session

s = Session(token_counter="gpt-4o")
s = Session(token_counter="meta-llama/Meta-Llama-3-8B-Instruct")

Custom counting callables are also accepted:

import tiktoken
from edgenote import Session

enc = tiktoken.encoding_for_model("gpt-4o")
s = Session(token_counter=enc.encode)

2. Automated Re-ranking

Automatically score and reorder retrieved documents at render time without manual relevance labels:

from edgenote import Session

s = Session(reranker="cross-encoder")
s.add("Quarterly financial filings")
s.add("Personnel roster and team structure")
s.add("Enterprise procurement guidelines")

result = s.render("What were the fourth-quarter operating expenditures?")

Pure-Python zero-overhead rankers are also available:

from edgenote import Session
from edgenote.reranker import BM25Reranker, TFIDFReranker

s = Session(reranker=BM25Reranker())
s = Session(reranker=TFIDFReranker())

3. Context Compression

Summarize low-priority context when token limits are reached instead of dropping text:

from edgenote import Session
from groq import Groq

client = Groq()

def summarize(text: str, max_tokens: int) -> str:
    response = client.chat.completions.create(
        model="llama3-8b-8192",
        messages=[
            {"role": "system", "content": f"Summarize in {max_tokens} tokens or fewer."},
            {"role": "user", "content": text},
        ],
        max_tokens=max_tokens,
    )
    return response.choices[0].message.content

s = Session(compressor=summarize)
s.add("Long technical specification...")
result = s.render("Summarize system latency limits", budget=2048)

4. Text Chunking

Split large documents before ingestion:

from edgenote.chunker import StructuralChunker, SemanticChunker

chunker = StructuralChunker(max_tokens=512, overlap_tokens=64)
chunks = chunker.chunk(document_text)

semantic_chunker = SemanticChunker(threshold=0.5)
semantic_chunks = semantic_chunker.chunk(document_text)

5. Framework Integrations

LangChain

from edgenote.integrations import from_langchain

result = from_langchain(
    docs,
    query="What is the operating budget?",
    pins=["Strictly cite sources using document headers."],
    budget=4000,
)
messages = result.to_messages()

LlamaIndex

from edgenote.integrations import from_llamaindex

result = from_llamaindex(
    nodes,
    query="Synthesize quarterly performance metrics.",
    pins=["Output format: Markdown table."],
)

Generic Dictionaries

from edgenote.integrations import from_dicts

docs = [
    {"text": "Annual recurring revenue reached $12M", "relevance": 0.95, "source": "Finance"},
    {"text": "Total headcount expanded to 120", "relevance": 0.40, "source": "HR"},
]
result = from_dicts(docs, query="Provide financial summary")

CLI & Model Cache

edgenote provides a command-line interface for verification and pre-caching neural weights:

# Verify installation and active backends
edgenote

# Run live prompt assembly demonstration
edgenote --demo

# Pre-cache weights for air-gapped environments
edgenote download bge-small

Author

Md Tareq Shah Alam

Release files for edgenote 0.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for edgenote 0.1.3
File Size Uploaded
edgenote-0.1.3.tar.gz 58.8 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for edgenote 0.1.3
File Interpreter ABI Platform
edgenote-0.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 58.8 MB

Release files / edgenote-0.1.3.tar.gz

Download URL edgenote-0.1.3.tar.gz
Size 58.8 MB
Tags Source
SHA-256 checksum
How to use checksums
ddda9d6b2fab3b33754c2d47b3490b4dcaa21498ef2c0ae2ba55c86901075418
BLAKE2b-256 checksum
How to use checksums
ba156e38338ace01e9dc2f33b9e5c0c349653ab6240255b87528fd38f96decff
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / edgenote-0.1.3-py3-none-any.whl

Download URL edgenote-0.1.3-py3-none-any.whl
Size 36.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
db0008050b0749014e93efeea9adf960146761d8444044259b1440025545cd60
BLAKE2b-256 checksum
How to use checksums
d7ad91107cdaf3f53acb21312f94e849164e83ef959542ffb2dde6a573f20444
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

This release

0.1.3 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page