Skip to main content

orcli -- open-refine client

A Python client library for interacting with OpenRefine via its REST API.

For simple project creation, data transformation, metadata management and export operations.

Find the orcli package on PyPI.

Features

  • Create and delete projects from local files.
  • Retrieve and manage project metadata.
  • Apply OpenRefine operations individually or in batches.
  • Load operations from JSON files.
  • Export project data using OpenRefine's supported export formats (TSV, CSV, JSON).
  • Retrieve column information and project models and convert row data to Python lists.
  • Error handling with detailed response logging.

Installation

Requirements

Requires Python 3.10+ and a running OpenRefine server instance.

Quick Start

  1. Download:

Clone the repository (see below).

git clone https://github.com/rkraasch/orcli.git
cd orcli

Or download from orcli from PyPI (see below).

python -m venv ./venv
./venv/bin/python -m pip install --upgrade pip
./venv/bin/python -m pip install orcli
./venv/bin/python -c "from orcli import Refine; print(Refine)"
  1. Optionally run tests:

Download and run OpenRefine, then run the tests as shown below.

python -m pytest tests/ -v

This creates temporary test projects in OpenRefine which should be cleaned up automatically. If a project prefixed pytest_orcli_ remains visible in the Open project tab (under #open-project), something went wrong.

  1. Basic usage:
from orcli import Refine

# Initialize the client
refine = Refine(base_url="http://127.0.0.1:3333")

# Create a project
project_id = refine.create_project("input_file.csv", "My Project")

# Get column names
columns = refine.get_column_names(project_id)
print(f"Columns: {columns}")

# Apply an operation
operation = {
    "op": "core/column-removal",
    "columnName": "unwanted_column",
    "description": "Remove column"
}
refine.apply_operation(operation, project_id)

# Export data
refine.export_data("output_file.tsv", fmt="tsv", project_id=project_id)

# Clean up
refine.delete_project(project_id)

Examples

The example projects can be found in ./examples/.

Example 1: Access Project Metadata

Demonstrates how to access and modify OpenRefine project metadata.

See files at ./examples/01_access-project-metadata/.

# Set metadata
refine.set_project_metadata("name", "Project Name", project_id)

# Get all projects
projects = refine.get_all_projects_metadata()
for pid, metadata in projects.items():
    print(f"{metadata['name']} (ID: {pid})")

# Find by name
project_id = refine.get_project_id_by_name("Project Name")

Example 2: Batch Operations

Apply multiple operations directly or load them from a JSON file.

See files at ./examples/02_batch-operations.

operations = [
    {"op": "core/column-removal", "columnName": "col1"},
    {"op": "core/column-removal", "columnName": "col2"}
]

refine.apply_operations(operations, project_id, wait=True)

refine.apply_operations_from_file(
    "operations.json",
    project_id,
    wait=True,
)

Example 3: Data Pipeline

Create a project, transform its data, export the result and clean up.

See files at ./examples/03_data-pipeline.

# Transform data
operations = [
    {
        "op": "core/column-removal",
        "columnName": "temp_field",
    },
    {
        "op": "core/text-transform",
        "engineConfig": {
            "facets": [],
            "mode": "row-based",
        },
        "columnName": "email",
        "expression": "value.toLowercase()",
        "onError": "keep-original",
        "repeat": False,
        "repeatCount": 10,
    },
]

refine.apply_operations(operations, project_id, wait=True)

Example 4: Batch Processing

Process every CSV file in a directory using the same operations file.

See files at ./examples/04_batch-processing.

for filename in os.listdir("input_dir/"):
    if filename.endswith(".csv"):
        project_id = refine.create_project(
            f"input_dir/{filename}",
            filename,
        )

        refine.apply_operations_from_file(
            "operations.json",
            project_id,
            wait=True,
        )

        refine.export_data(
            f"output_dir/{filename}",
            fmt="csv",
            project_id=project_id,
        )

        refine.delete_project(project_id)

Example 5: Logging

Configure the Python logger to control the output produced by orcli at different log levels.

See files at ./examples/05_logging/.

import logging

from orcli import Refine

logger = logging.getLogger("orcli.client")

handler = logging.StreamHandler()
handler.setFormatter(
    logging.Formatter("%(levelname)s: %(message)s")
)

logger.addHandler(handler)
logger.propagate = False

# DEBUG: show all diagnostic messages.
logger.setLevel(logging.DEBUG)

refine = Refine()

# Use Refine as usual.
project_id = refine.create_project(
    project_file="input.csv",
    project_name="Logging Example",
)

# INFO: suppress DEBUG messages.
logger.setLevel(logging.INFO)

# WARNING: suppress DEBUG and INFO messages.
logger.setLevel(logging.WARNING)

# ERROR: only show errors and critical messages.
logger.setLevel(logging.ERROR)

API Reference

Initialization

Refine(base_url=None)

Parameters:

Methods

Method Description
create_project(file, name) Create project
delete_project(project_id) Delete project
get_all_projects_metadata() Get all projects
get_project_id_by_name(name) Find project by name
set_project_metadata(field, value, id) Update metadata
apply_operation(op, id, wait) Apply operation
apply_operations(ops, id, wait) Apply multiple
apply_operations_from_file(file, id, wait) Load from file
get_models(id) Get models
get_column_names(id) Get columns
export_data(file, fmt, id) Export data
rows_as_list(data) Convert rows
wait_until_idle(id, delay) Wait for completion

Configuration

Logging

orcli uses Python's standard logging module. The library does not configure logging output itself, allowing applications to decide which log messages to display.

See Example 5: Logging for a code example.

Custom server

refine = Refine(base_url="http://example.com:3333")

Troubleshooting

  • ConnectionError: Ensure OpenRefine is running, verify the server URL and check network connectivity.
  • CSRF token errors: Check server status and logs.
  • FileNotFoundError: Use absolute paths, verify the file exists and check permissions.

Support and Contribute

Open an issue on the project's issues tab on Github.

Or contribute via Github:

  • fork the repository,
  • create a feature branch,
  • commit changes,
  • push to branch,
  • open Pull Request.

References

References:

Similar Projects:

License

This project is released under CC0 1.0 Universal License.

For a plain text version see this project's LICENSE file or visit creativecommons.org.

Release files for orcli 0.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for orcli 0.1.3
File Size Uploaded
orcli-0.1.3.tar.gz 12.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for orcli 0.1.3
File Interpreter ABI Platform
orcli-0.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 22.3 kB

Release files / orcli-0.1.3.tar.gz

Download URL orcli-0.1.3.tar.gz
Size 12.3 kB
Tags Source
SHA-256 checksum
How to use checksums
f8ea9f3d5ecaf5ce6e9395850610cebca415ddabe6e9db9a82c5f6b72baee87d
BLAKE2b-256 checksum
How to use checksums
532fe590a228d763d4e7fed0ce629b9bd5c54a7be0cfd0c683b69eef471431d9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release files / orcli-0.1.3-py3-none-any.whl

Download URL orcli-0.1.3-py3-none-any.whl
Size 10.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1c992d0ed3b1cb995826d46473f3a708d9786cd815e7f7c5245aae2c351c4eb9
BLAKE2b-256 checksum
How to use checksums
6ce44caa99dfe8f03584f5b659967f38842ef5023a2b0f19f9c96fe3f0b9c786
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.7

Release history Release notifications | RSS feed

This release

0.1.3 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page