Skip to main content
LogLens AI

Detect anomalies by meaning. Explain why they matter. Group them into incidents. Monitor services in real time.

PyPI · Documentation · Website · Docker Hub · Benchmarks

PyPI version Python versions CI status MIT License

PyPI downloads per month Total PyPI downloads Docker pulls Docker image size GitHub stars Repository views

Overview

Modern applications can generate thousands or millions of log lines. Finding the few lines that actually represent a failure, degradation, or security event is often more difficult than generating the logs themselves.

LogLens AI approaches this problem as an anomaly detection and incident analysis pipeline.

Instead of treating every log line independently, it:

  1. Parses and normalizes incoming logs.
  2. Learns recurring log templates.
  3. Represents templates using statistical or semantic embeddings.
  4. Detects unusual behavior.
  5. Groups related anomalies into incident families.
  6. Explains why each anomaly was detected.
  7. Optionally generates an AI-assisted root-cause narrative.
  8. Delivers the result through the CLI, live monitoring, alerts, or offline HTML reports.

The detection pipeline runs locally, making the project suitable for environments where logs should remain on the machine or infrastructure that produced them.

Table of Contents

Key Features

Semantic anomaly detection

Detect unusual log behavior using the semantic characteristics of log messages rather than relying exclusively on fixed keywords or hand-written rules.

Explainable results

Every detected anomaly includes human-readable reasons based on factors such as:

  • Severity
  • Template rarity
  • Frequency changes
  • Semantic distance
  • Burst behavior
  • Historical patterns

The goal is not simply to say "this line is anomalous", but to provide context for why it was flagged.

Incident grouping

Repeated anomalies can represent a single underlying incident.

LogLens AI groups related events into incident families so that hundreds of repeated errors can be represented as one actionable incident rather than hundreds of individual alerts.

Real-time monitoring

Monitor running services directly from the command line.

Supported sources include:

  • Docker logs
  • Kubernetes logs
  • journald
  • Custom streaming commands

The live watcher focuses on anomalous events instead of forcing developers to manually scan an entire log stream.

Application self-alerting

Applications can integrate LogLens AI directly through the Python SDK.

The integration can surface serious events through:

  • Slack
  • Microsoft Teams
  • Email

Alerting is designed to be asynchronous, rate-limited, and de-duplicated so that the monitoring layer does not become a reliability risk for the application itself.

Optional AI root-cause analysis

LogLens AI can optionally use a user-provided LLM API key to generate higher-level incident explanations.

Supported providers include:

  • OpenAI
  • Azure
  • Groq

The RCA layer operates on grouped anomaly summaries rather than transmitting the complete log stream.

Offline HTML reports

Generate self-contained HTML reports containing:

  • Incident summaries
  • Severity breakdowns
  • Service-level information
  • Score distributions
  • Detected anomalies
  • Optional RCA narratives

Reports can be viewed without a cloud dashboard or external web service.

Multiple detection engines

LogLens AI provides three detection modes:

Mode Approach Primary goal
fast Statistical / TF-IDF Fast general-purpose detection
turbo Optimized statistical pipeline Higher throughput
deep Transformer-based semantic embeddings Deeper semantic analysis

The core detection pipeline does not require an external AI service.


Installation

Python

Requires Python 3.10+.

Install the standard package with:

pip install loglensai

For transformer-based semantic detection:

pip install "loglensai[deep]"

Docker

LogLens AI also provides multi-architecture Docker images for amd64 and arm64.

Example:

docker run --rm -v "$PWD:/data" loglensai/loglens analyze --source app.log

Quick Start

# Analyze a log file
loglens analyze --source app.log

# Use the optimized high-throughput detector
loglens analyze --source app.log --turbo

# Use semantic transformer-based detection
loglens analyze --source app.log --deep

# Generate an offline incident report with optional RCA
loglens analyze --source app.log --turbo --rca --html report.html

# Monitor a running Docker service
loglens watch "docker logs -f my-api"

# Ask a natural-language question about detected anomalies
loglens ask "why did the payment service start timing out?" --source app.log

# Run the reproducible benchmark
loglens benchmark labeled.log --min-f1 0.90

Python SDK

LogLens AI can also be integrated directly into Python applications.

For eg:-

from loglens import analyze
result = analyze("app.log")                 # or lines=[...], cmd="docker logs api"
for a in result.anomalies:
    print(a.level, a.score, a.message, a.reasons)
print(result.rca().report)                  # AI root-cause (BYO key)
  • analyze() / analyze_async() - "here are logs, give me the problems."
  • LogLensHandler - drop into Python's logging so your app raises its own alarm.
  • LiveDetector - feed a custom stream line-by-line, get anomalies out (powers watch).
  • .rca(), .ask(...), .save_html(...), .save_rca(...) on any result or live session.

The SDK provides interfaces for:

  • Batch log analysis
  • Asynchronous analysis
  • Python logging integration
  • Live stream detection
  • Root-cause analysis
  • Natural-language questions
  • HTML report generation
  • Alerting

For example, an application can initialize the monitoring layer with:

loglens.init(app_name="checkout-api")

This allows LogLens AI to monitor application events and surface serious anomalies without requiring a separate logging agent.

See the SDK documentation for the complete API reference.


How It Works

LogLens AI processes logs through several stages.

1. Parse

The input stream is parsed using an automatically detected log format.

Supported formats include Apache, Linux, Mac, HDFS, Spark, Zookeeper, OpenStack, Thunderbird, BGL, HealthApp, and generic logs.

2. Template extraction

Similar log messages are converted into reusable templates.

For example, multiple messages such as:

Connection failed for user 18372

and

Connection failed for user 92451

can be represented by a common structural template rather than treated as completely unrelated events.

3. Representation

The detection engine represents log templates using either:

  • TF-IDF-based representations for fast and turbo
  • Transformer embeddings for deep

Semantic representations are generated at the template level where possible, reducing unnecessary computation on repeated messages.

4. Anomaly detection

The detection system combines multiple signals, including:

  • Severity
  • Template rarity
  • Embedding distance
  • Frequency
  • Chronic-pattern damping
  • Global rarity

These signals are combined into a continuous anomaly score.

5. Incident grouping

Related anomaly events are grouped into incident families.

This reduces alert noise and provides a higher-level view of what is happening within the system.

6. Explanation

Each anomaly is accompanied by human-readable reasoning describing the signals that contributed to the detection.

7. Delivery

Results can be delivered through:

  • CLI output
  • Live monitoring
  • Slack
  • Microsoft Teams
  • Email
  • Offline HTML reports
  • Python SDK

Optional LLM-based RCA can add a higher-level narrative to the detected incident.


Benchmarks

LogLens AI includes a reproducible benchmarking harness based on labeled datasets from Loghub.

BGL Dataset

The current benchmark evaluates 500,000 BGL log lines containing 206,847 labeled alerts.

Mode Engine Precision Recall F1 Approx. throughput
fast Statistical 0.901 1.000 0.948 ~6,700 lines/s
turbo Optimized statistical 0.901 1.000 0.948 ~7,300 lines/s
deep Semantic embeddings 0.917 1.000 0.957 ~3,400 lines/s

These figures are benchmark results on the specified dataset and configuration; they should not be interpreted as universal performance guarantees.

Reproduce it yourself (don't take our word for it):

loglens benchmark path/to/BGL.log --format bgl --supervised

On the bundled 2,000-line sample (benchmarks/BGL_2k.log), the supervised head scores F1 0.932 (precision 0.902, recall 0.965) and the unsupervised path reaches 1.000 recall — see BENCHMARKS.md for the full measured baseline (accuracy, throughput, memory).

The complete methodology and reproduction instructions are available in BENCHMARKS.md.

Generalization Test

A separate evaluation uses 500,000 normal lines from the Sandia Thunderbird dataset without retuning the detector.

Mode False-alarm rate Specificity Approx. throughput
fast 0.68% 99.32% ~8,600 lines/s
deep 0.67% 99.33% ~1,800 lines/s

Injected Incident Test

The project also includes a synthetic "needle-in-a-haystack" evaluation involving 30 injected incidents across six log formats.

The tested incidents include scenarios such as:

  • Kernel panic
  • Out-of-memory conditions
  • Disk failures
  • Security events
  • Data corruption

The current evaluation detected all 30 injected incidents.

For methodology and reproduction details, see BENCHMARK.md.


Privacy and Data Handling

LogLens AI is designed around a local-first architecture.

The core detection process does not require logs to be uploaded to a cloud observability platform.

Optional AI functionality requires an external LLM provider and is therefore subject to that provider's network and data-handling policies.

When RCA or natural-language analysis is enabled, LogLens AI sends grouped anomaly summaries rather than the complete raw log stream.

Alerting can also operate without an LLM using the built-in detection and explanation mechanisms.


Docker

The project publishes multi-architecture images supporting:

  • Linux amd64
  • Linux arm64

Analyze a mounted log file:

docker run --rm -v "$PWD:/data" loglensai/loglens analyze --source app.log

For live Docker monitoring, the Docker socket can be mounted read-only and used as the source for the watcher.

Available image variants include the standard release and an optional image containing the neural detection dependencies.


Project Structure

The project is organized around several major components:

Detection engine Parsing, template extraction, feature generation, anomaly scoring, and incident grouping.

CLI Commands for analysis, monitoring, benchmarking, reporting, and natural-language queries.

Python SDK Programmatic integration for applications and custom pipelines.

Live monitoring Streaming detection for Docker, Kubernetes, journald, and custom commands.

Alerting Asynchronous notification delivery with deduplication and rate limiting.

AI layer Optional root-cause analysis and natural-language interaction.

Reporting Offline HTML reports and benchmark outputs.


Roadmap

Planned improvements include:

  • Improved chronic-noise handling for Linux and macOS daemon logs
  • A prebuilt GitHub Action for CI log analysis
  • Additional alert integrations such as PagerDuty and Opsgenie
  • Generic webhook support
  • A lightweight optional web interface for generated reports

Reproducibility

Benchmark results are intended to be reproducible.

The repository includes the evaluation harness and benchmark documentation needed to run the supported experiments independently.

See:


Contributing

Contributions are welcome.

You can contribute through:

  • Bug reports
  • Feature requests
  • Documentation improvements
  • Benchmark improvements
  • New log format support
  • Detection algorithms
  • Integrations
  • Pull requests

Please open an issue before starting a major architectural change so the proposed direction can be discussed.

License

LogLens AI is released under the MIT License.

See LICENSE for details.


LogLens AI

Understand your logs. Find the incident. Fix the problem.

Metadata

Release files for loglensai 0.9.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for loglensai 0.9.0
File Size Uploaded
loglensai-0.9.0.tar.gz 2.0 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for loglensai 0.9.0
File Interpreter ABI Platform
loglensai-0.9.0-py3-none-any.whl Python 3 none any Details

Total release size: 3.1 MB

Release files / loglensai-0.9.0.tar.gz

Download URL loglensai-0.9.0.tar.gz
Size 2.0 MB
Tags Source
SHA-256 checksum
How to use checksums
e9050481bb45115e0378829d5ea4120d4c8eac5c3fad98eaae0a4c8f653d63b2
BLAKE2b-256 checksum
How to use checksums
191880ef22a0cb11ec89e61a8dfe8af000435fa45264c3a7915b27e6683f5264
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release files / loglensai-0.9.0-py3-none-any.whl

Download URL loglensai-0.9.0-py3-none-any.whl
Size 1.1 MB
Tags Python 3
SHA-256 checksum
How to use checksums
96aa93dbd60c8e4476edadea2411cb872e6e3e69e267a0f269c1cb1ccea91686
BLAKE2b-256 checksum
How to use checksums
0e0d3bbc84e406c5fe5e99378f0819f300ad8aae94906ac7e95958f492b61d91
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release history Release notifications | RSS feed

This release

0.9.0 This release

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page