Skip to main content

API client for metadata curation platforms

Project description

Metadata Curation Client

API client for external partners to integrate with metadata curation platforms.

Installation

pip install metadata-curation-client

Basic Usage

from metadata_curation_client import CurationAPIClient, PropertyType

# Initialize client
client = CurationAPIClient("http://localhost:8000")

# Create source
source = client.create_source({
    "name": "My Archive",
    "description": "Digital editions from our collection"
})

# Create controlled vocabulary property
language_prop = client.create_property({
    "technical_name": "language",
    "name": "Language", 
    "type": PropertyType.CONTROLLED_VOCABULARY,
    "property_options": [{"name": "English"}, {"name": "German"}]
})

# Create free text property
description_prop = client.create_property({
    "technical_name": "description",
    "name": "Description", 
    "type": PropertyType.FREE_TEXT
})

# Create edition
edition = client.create_edition({
    "source_id": source["id"],
    "source_internal_id": "my_001"
})

# Create properties for each type
genre_prop = client.create_property({
    "technical_name": "genre",
    "name": "Genre", 
    "type": PropertyType.CONTROLLED_VOCABULARY,
    "property_options": [
        {"name": "Poetry"}, {"name": "Prose"}, {"name": "Drama"}
    ]
})

has_annotations_prop = client.create_property({
    "technical_name": "has_annotations",
    "name": "Has Annotations", 
    "type": PropertyType.BINARY
})

year_prop = client.create_property({
    "technical_name": "publication_year",
    "name": "Publication Year", 
    "type": PropertyType.NUMERICAL
})

description_prop = client.create_property({
    "technical_name": "description",
    "name": "Description", 
    "type": PropertyType.FREE_TEXT
})

# Example 1: CONTROLLED_VOCABULARY suggestion
# First get the property option ID
properties = client.get_properties()
genre_prop = next(p for p in properties if p["technical_name"] == "genre")
poetry_option = next(opt for opt in genre_prop["property_options"] if opt["name"] == "Poetry")

client.create_suggestion({
    "source_id": source["id"],
    "edition_id": edition["id"],
    "property_id": genre_prop["id"],
    "property_option_id": poetry_option["id"]
})

# Example 2: BINARY suggestion (uses property_option_id)
# Binary properties always have options with ID 1 (true/1) and ID 2 (false/0)
# Get the "true" option (usually ID 1)
binary_props = client.get_properties()
has_annotations_prop = next(p for p in binary_props if p["technical_name"] == "has_annotations")
true_option = next(opt for opt in has_annotations_prop["property_options"] if opt["name"] == "1")

client.create_suggestion({
    "source_id": source["id"],
    "edition_id": edition["id"],
    "property_id": has_annotations_prop["id"],
    "property_option_id": true_option["id"]  # For "yes"/"true" value
})

# Example 3: NUMERICAL suggestion (uses custom_value)
client.create_suggestion({
    "source_id": source["id"],
    "edition_id": edition["id"],
    "property_id": year_prop["id"],
    "custom_value": "2025"  # Note: numerical values are sent as strings
})

# Example 4: FREE_TEXT suggestion (uses custom_value)
client.create_suggestion({
    "source_id": source["id"],
    "edition_id": edition["id"],
    "property_id": description_prop["id"],
    "custom_value": "This is a detailed description of the edition."
})

# Mark ingestion complete
client.mark_ingestion_complete(source["id"])

Property Types

  • PropertyType.CONTROLLED_VOCABULARY - Predefined options
  • PropertyType.FREE_TEXT - Free text
  • PropertyType.BINARY - True/false values
  • PropertyType.NUMERICAL - Numeric values

API Reference

See the docstrings in curation_api_client.py for detailed method documentation.

Enhanced Integration with SourceManager

For more sophisticated integrations, we also provide a higher-level abstraction in source_manager.py that mirrors some of the conveniences of our internal extractors:

from metadata_curation_client import CurationAPIClient, PropertyType, SourceManager, PropertyBuilder

# Initialize client and create source
client = CurationAPIClient("http://localhost:8000")
source = client.get_source_by_technical_name("my_data_source")
if not source:
    source = client.create_source({
        "name": "My Data Source",
        "description": "My collection of digital editions",
        "technical_name": "my_data_source"
    })

# Define properties using helper builders
property_definitions = [
    PropertyBuilder.controlled_vocabulary(
        "example_genre", "Genre", ["Poetry", "Prose", "Drama"]
    ),
    PropertyBuilder.binary(
        "example_has_annotations", "Has Annotations"
    ),
    PropertyBuilder.numerical(
        "example_year", "Publication Year"
    )
]

# Initialize the source manager - this will:
# - Fetch all existing data
# - Build lookup tables
# - Create any missing properties
manager = SourceManager(client, source['id'], property_definitions)

# Efficiently get or create edition using lookup tables
edition = manager.get_or_create_edition("book_001")

# Create suggestions in a batch with validation and deduplication
manager.create_suggestions_batch(
    edition["id"],
    {
        "example_genre": "Poetry",
        "example_has_annotations": True,
        "example_year": 2022
    }
)

# Mark ingestion complete (updates timestamp)
manager.finish_ingestion()

Benefits of the SourceManager

The SourceManager provides several advantages for more complex integrations:

  1. Reduced API Calls: Prefetches data to minimize API requests
  2. Lookup Tables: Maintains efficient in-memory lookups for editions, properties, and suggestions
  3. Automatic Property Creation: Creates properties from definitions as needed
  4. Validation: Automatically validates values based on property types
  5. Deduplication: Avoids creating duplicate suggestions
  6. Builder Helpers: Provides convenient builder classes for creating properties and sources
  7. Timestamp Management: Automatically updates the last ingestion timestamp

For a complete example, see example_with_source_manager.py.

Choosing the Right Approach

  • Basic API Client: For simple integrations or when you need complete control over the process
  • SourceManager: For more complex integrations where efficiency and convenience are priorities

Both approaches use the same underlying API endpoints and data models, so you can choose the one that best fits your needs or even mix them as required.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

metadata_curation_client-0.3.0.tar.gz (11.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

metadata_curation_client-0.3.0-py3-none-any.whl (11.2 kB view details)

Uploaded Python 3

File details

Details for the file metadata_curation_client-0.3.0.tar.gz.

File metadata

  • Download URL: metadata_curation_client-0.3.0.tar.gz
  • Upload date:
  • Size: 11.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.1

File hashes

Hashes for metadata_curation_client-0.3.0.tar.gz
Algorithm Hash digest
SHA256 0a66436e2c52fd998dc5e06d5757b088fc8bd73d37c7f441e4ae7332ee05c736
MD5 6e27a048a35f54baf31354e43a212617
BLAKE2b-256 ef77df9ef0039745cac0598a9a78c8ec5dd950caba2aaaad52ac55b5ee810309

See more details on using hashes here.

File details

Details for the file metadata_curation_client-0.3.0-py3-none-any.whl.

File metadata

File hashes

Hashes for metadata_curation_client-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e23de84339c59931e8a2cb0a458e3f928820e8f22862eee3876213b6533e4f75
MD5 d45acaa2cfc0fb11ae0f0c1fab296726
BLAKE2b-256 eed083735df5fb003994042e0e274bdd0a49ecc9a85993fb6820753a91594bce

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page