ojs
Tools for working with the Open Journal Systems (OJS) API.
Pulls submissions, publications, reviews, users, and publication view statistics
from an OJS journal's /api/v1/* REST API, downloads the attached file artifacts
(manuscripts, revisions, reviewer attachments, and production galleys), and
normalizes the JSON into typed relational tables backed by polars. Incremental
sync re-pulls only what changed since the last run, so routine top-ups stay
cheap. A typed schema layer (Column/Table classes) is the single source of
truth for normalization and doubles as exportable column documentation. Also
normalizes the OJS dashboard's Articles and Reviews CSV report exports. Ships a
Typer CLI for the common fetch, download, and normalize workflows. Built against
the OJS 3.3 REST API; other versions
are untested and may differ, as the REST API saw breaking changes between 3.2
and 3.3.
Project Structure
ojs/
├── cli.py # Typer CLI: init, articles, reviews, api (+ schema docs)
├── errors.py # OjsError hierarchy raised by the run entry points
├── paths.py # Output directory resolution (argument → env var → default)
├── schema.py # Typed schema framework: Column/Table, apply(), doc export
├── utils.py # HTML stripping + localized-field extraction
├── website/ # Website CSV-export pipelines
│ ├── run.py # run_report_fetch / run_norm entry points
│ ├── reports.py # Authenticated report-CSV fetch (OJS login + download)
│ ├── articles/ # Wide CSV → submissions, authors, editors, decisions
│ │ ├── normalize.py # Unpivot the wide CSV into the four tables
│ │ └── schemas.py # Submissions/Authors/Editors/Decisions schemas
│ └── reviews/ # Long CSV → reviews
│ ├── normalize.py # Rename and type-cast review data
│ └── schemas.py # Reviews schema
└── api/ # REST pipeline
├── run.py # run_fetch / run_download / run_norm entry points
├── client.py # OJS REST client (httpx, pagination, retry, early-stop)
├── files.py # Submission file artifact downloads (disk layout, manifest)
├── normalize.py # JSON → relational tables (schema-driven)
├── schemas.py # API table schema classes
├── sync.py # Incremental sync: high-water-mark state, raw-JSON upsert
└── swagger.json # OJS API reference (snapshot)
The CLI is a thin shell over api/run.py and website/run.py: every command
resolves its options, calls the matching run_* function, and turns an error
into an exit code. The pipelines themselves are importable — see
Using ojs as a library.
Installation
uv tool install ojs
As a project dependency:
uv add ojs
From GitHub instead of PyPI:
uv tool install git+https://github.com/gitronald/ojs.git
# or, as a dependency: uv add git+https://github.com/gitronald/ojs.git
From source (for development):
git clone https://github.com/gitronald/ojs.git
cd ojs
uv sync
Configuration
The CLI reads from a .env file in the current directory. Run ojs init to
scaffold one — it prompts for the journal URL and API token, and writes .env
with 0600 permissions:
ojs init
Getting an API key. In OJS, open your user profile
(https://example.org/index.php/myjournal/user/profile), select the API Key
tab, check Enable external applications with the API key to access this
account, and copy the key — use the (re)generate button if one isn't set
yet.
Values can also come from the environment. A user-level config file is loaded as
a fallback for anything not set in the current directory's .env (which takes
precedence): ~/.config/ojs/.env by default, or the file named by
OJS_CONFIG_PATH.
| Variable | Default | Purpose |
|---|---|---|
OJS_CONFIG_PATH |
~/.config/ojs/.env |
User-level config file, loaded as a fallback for values the current directory's .env and the environment don't set |
OJS_BASE_URL |
(required for api) |
OJS journal URL (e.g. https://example.org/index.php/myjournal) |
OJS_API_KEY |
(required for api) |
OJS API token |
OJS_USERNAME |
(required for website fetch) |
Editorial-manager login username |
OJS_PASSWORD |
(required for website fetch) |
Password for that account |
OJS_REVIEWS_REPORT_URL |
(required for reviews fetch) |
Review Report export URL (absolute, or a path relative to OJS_BASE_URL) |
OJS_ARTICLES_REPORT_URL |
(required for articles fetch) |
Articles Report export URL (absolute, or a path relative to OJS_BASE_URL) |
OJS_DOWNLOADS_DIR |
data/ojs-website |
Website CSV exports + normalized output |
OJS_ARTICLES_DIR |
$OJS_DOWNLOADS_DIR/articles |
Articles output dir |
OJS_REVIEWS_DIR |
$OJS_DOWNLOADS_DIR/reviews |
Reviews output dir |
OJS_API_DIR |
data/ojs-api |
API JSON dump dir |
OJS_FILES_DIR |
$OJS_API_DIR/files |
Where downloaded submission files land |
CLI Commands
norm reads the typed schema classes directly — no separate step is required.
schema exports a table_schemas.csv documenting each table's columns, dtypes,
source mapping, and whether each column appears in the normalized output
(in_output).
API
Fetch raw JSON from the REST API, download file artifacts, and normalize into relational tables.
ojs api fetch # fetch raw JSON from the OJS REST API
ojs api download # download submission file artifacts (PDFs, etc.)
ojs api norm # normalize API JSON into relational tables
ojs api schema # export table_schemas.csv docs
Articles
Fetch and normalize the OJS dashboard's Articles Report CSV export.
ojs articles fetch # download the latest articles export from the website
ojs articles norm # normalize the most recent articles export
ojs articles schema # export table_schemas.csv docs
Reviews
Fetch and normalize the OJS dashboard's Review Report CSV export.
ojs reviews fetch # download the latest reviews export from the website
ojs reviews norm # normalize the most recent reviews export
ojs reviews schema # export table_schemas.csv docs
Website report fetch
The Articles and Review reports are OJS dashboard exports (Statistics > Tools),
not REST endpoints, so fetch authenticates differently from the api commands:
it logs in with OJS_USERNAME / OJS_PASSWORD to establish a session, then
downloads the report at OJS_REVIEWS_REPORT_URL / OJS_ARTICLES_REPORT_URL. The
CSV lands in OJS_DOWNLOADS_DIR as reviews-<YYYYMMDD>.csv /
articles-<YYYYMMDD>.csv, where the matching norm command picks up the newest
file. The report URLs are instance-specific — set them in .env (see
.env.example).
Article view stats
ojs api fetch also pulls publication view stats from the OJS /stats/publications/*
endpoints (skip with --no-stats). The API only exposes aggregated counts — the
finest granularity is daily (there are no per-event timestamps).
| Flag | Default | Purpose |
|---|---|---|
--stats / --no-stats |
on | Toggle stats collection (e.g. when the API key lacks stats access) |
--stats-interval |
day |
Timeline granularity: day or month |
--stats-since |
(none) | dateStart filter (YYYY-MM-DD) |
--stats-until |
(none) | dateEnd filter (YYYY-MM-DD) |
ojs api norm then writes three extra tables:
publication_stats— one row per published submission with abstract, all-galley, PDF, HTML, and other view totals.views_timeline— long format (submission_id,date,interval,views,kind) with a per-submission abstract and galley series.intervalrecords the granularity (dayormonth) a point was fetched at, so a file mixing both stays separable — filter on it rather than summing across intervals.views_timeline_totals— long format (date,interval,views,kind) with the journal-wide abstract and galley series, from the aggregate/stats/publications/{abstract,galley}endpoints (the data behind the OJS statistics-page graph). Use this for journal-wide totals rather than summingviews_timeline.
If the API key lacks stats access, fetch prints a warning and skips the stats files, and norm simply omits the two tables.
Submission files
OJS attaches the actual file artifacts (manuscripts, revisions, reviewer
attachments, production galleys) to each submission. ojs api download fetches
their metadata and then downloads the binaries.
ojs api fetch --files # also dump file metadata -> submission_files.json
ojs api download # download all files for all submissions
ojs api download -s 123 -s 456 # only these submissions (repeatable)
ojs api download --type galleys # only published galley files
ojs api download --type review # only review files / revisions / attachments
ojs api download --file-stage 4 --file-stage 15 # raw fileStage ids
ojs api download --no-revisions # current files only, skip prior revisions
| Flag | Default | Purpose |
|---|---|---|
--submission-id / -s |
all | Limit to these submission ids (repeatable) |
--type |
all |
all, galleys (published), or review |
--file-stage |
(none) | Raw fileStage id(s); overrides --type |
--revisions / --no-revisions |
on | Also download prior revisions of each file |
--fetch / --no-fetch |
on | Refresh file metadata first (off: use stored JSON) |
Files are laid out under OJS_FILES_DIR as
<submission_id>/<stage>/<fileId>_<name>. A manifest (manifest.json) records
every artifact by its immutable physical fileId, so reruns skip files already
on disk — new uploads and revisions are downloaded incrementally.
Rounds and revisions. OJS tracks two distinct axes. A file's stage
(fileStage) says where in the workflow it lives; review files additionally
carry an assocId naming the review round they belong to. Separately, each
file's revisions[] holds prior uploads of that same logical file. ojs api norm writes a submission_files table with one row per current file, including
file_stage_label, review_round_id (joins to review_assignments.round_id),
and revision_count. Downloads cover the current file plus every revision, each
keyed by its own fileId.
Downloading files requires an API token with permission to view them; the API
returns 403 for files the key cannot access.
Incremental fetch
By default ojs api fetch does a full cold pull. For routine top-ups, --incremental fetches only what changed since the last successful sync and merges it into the existing JSON dumps, so ojs api norm stays a stateless re-derivation from the complete files.
| Flag | Purpose |
|---|---|
--incremental / -i |
Fetch only records changed since the last sync, merging into the JSON dumps |
--since YYYY-MM-DD |
Override the stored watermark (implies --incremental) |
--full |
Force a complete pull and reset the sync state |
How it works:
- A high-water mark lives in
data/ojs-api/sync_state.json(the last sync time, plus each submission'sdateLastActivity). It advances only after a run fully succeeds, so a failed fetch never skips records on the next run. - Submissions and extended submissions are pulled newest-first by
dateLastActivityand stop early at the watermark. Publication details are skipped for submissions whosedateLastActivityis unchanged — the biggest saving, since that endpoint costs one request per submission. - A one-day overlap buffer re-pulls the boundary on each run; merges are idempotent (upsert by id), so the overlap is harmless.
- View stats:
publication_stats(cumulative totals) is always pulled in full, while the dailyviews_timelineis re-pulled over a rolling window and merged by(submission_id, interval, date, kind), refreshing recent buckets without dropping history. - Users are always pulled in full — the API exposes no recency sort for users.
The OJS API has no server-side "modified since" filter, so incremental cannot detect upstream deletions; run ojs api fetch --full periodically to reconcile.
Upgrade note. The view-stats timelines gained an interval column (day vs
month) that is part of the incremental merge key. Timelines written before that
change lack it, so the first --incremental run after upgrading would leave the
legacy points sitting alongside the re-fetched ones, since their keys differ. Run
a one-time ojs api fetch --full after upgrading to rewrite
views_timeline.json and views_timeline_totals.json cleanly.
Normalized output
Every norm command writes one CSV per table, named for the table, next to a
table_schemas.csv written by the matching schema command:
| Command | Output directory |
|---|---|
ojs api norm |
$OJS_API_DIR/normalized/ (default data/ojs-api/normalized/) |
ojs articles norm |
$OJS_ARTICLES_DIR (default data/ojs-website/articles/) |
ojs reviews norm |
$OJS_REVIEWS_DIR (default data/ojs-website/reviews/) |
ojs api norm — eight tables from the REST API:
| Table | One row per |
|---|---|
submissions |
Submission — status, stage, dates (incl. last activity), latest review round, type, DOI, and a first-author summary |
publications |
Publication version — title, abstract, issue, pages, license, galley count |
authors |
Author per submission, author_number in display order, joined to OJS accounts via user_id |
review_assignments |
Reviewer assignment — round, status, response and review due dates |
submission_files |
Current file artifact — stage, review round, revision count, uploader, URL |
publication_stats |
Published submission — abstract, galley, PDF, HTML, and other view totals |
views_timeline |
Submission/date/kind view count (long format) |
views_timeline_totals |
Journal-wide date/kind view count (long format) |
The last three appear only when stats were fetched, and submission_files only
after ojs api fetch --files or ojs api download.
ojs articles norm — four tables unpivoted from the wide Articles Report:
submissions (one row per submission, including the language, rights, subjects,
and disciplines metadata the REST API doesn't expose), plus authors, editors,
and decisions, each one row per numbered (Author N) / (Editor N) /
decision column group.
ojs reviews norm — a single reviews table, one row per reviewer
assignment, carrying the full date chain (assigned, notified, confirmed,
completed, acknowledged, reminded), the response and review overdue day counts,
the recommendation, and the reviewer's comments.
Using ojs as a library
Every pipeline the CLI runs is an importable function, so the package can be
driven in-process instead of through subprocess:
from ojs.api.run import run_fetch, run_norm
fetch = run_fetch(incremental=True, stats=False)
print(fetch.submissions.fetched, "changed;", fetch.submissions.total, "on disk")
norm = run_norm()
print(norm.rows) # {"submissions": 354, "publications": 361, ...}
The website pipelines mirror this, parameterized by report. They have their own
run_norm, so import the modules rather than the functions when a caller drives
both pipelines:
from ojs.website import run as website
export = website.run_report_fetch("reviews") # -> Path to the downloaded CSV
result = website.run_norm("reviews", input_file=export)
Three conventions make these usable from another codebase:
- Arguments before environment. Apart from the website entry points' leading
report, every argument is keyword-only, and the ones naming a location or a credential default toNone, meaning "read the environment" — the same defaults the CLI uses. Passbase_url,api_key,out_dir, and friends explicitly to bypass.enventirely.ojs.pathsexposes the same directory resolution (api_dir(),articles_dir(), …) so a caller can ask where output lands rather than reconstructing the defaults. An override reaches the directories derived from it, sorun_norm("articles", downloads_dir="custom")writes its tables tocustom/articlesrather than to$OJS_ARTICLES_DIR— likewisereviews_dirfromdownloads_dir, andfiles_dirfromapi_dir. The rule is one rule in both directions: anything the caller passes beats anything in the environment, and theOJS_*_DIRvars apply whenever no directory argument is given (as in every CLI invocation). - Results, not printed lines.
run_fetchreturns aFetchResult(per-datasetfetched/totalcounts, whether stats succeeded, whether the run was incremental, the output directory);run_normreturns the table names and row counts;run_downloadreturns the downloaded and failed records. - Exceptions, not exit codes. Library code raises
ojs.errors.OjsError—ConfigError(missing credentials or report settings),OptionError(contradictory or malformed arguments),MissingDataError(a required input is not on disk). The CLI catches these and maps them back to its usual messages and exit codes.
Progress goes to the ojs logger, which the package fits with a NullHandler,
so an embedding caller sees nothing on the console by default. To get the same
lines the CLI prints, attach a handler on stdout (basicConfig alone would send
them to stderr):
import logging
import sys
handler = logging.StreamHandler(sys.stdout)
handler.setFormatter(logging.Formatter("%(message)s"))
logger = logging.getLogger("ojs")
logger.addHandler(handler)
logger.setLevel(logging.INFO)
Security & privacy
- The API token lives in
.env(theinitprompt hides input), alongsideOJS_PASSWORDfor the websitefetchcommands..envis gitignored and written0600— keep it out of version control and out of shared locations. - The API JSON dumps contain personal data pulled from OJS:
users.jsonholds user records including email addresses, and the author/submission tables carry author names, emails, and ORCIDs. These files are written with the process umask (typically0644, i.e. world-readable). On a shared or multi-user host, run with a restrictive umask (e.g.umask 077) or pointOJS_API_DIRat a private directory so other local users can't read them. - The fetched website reports (
reviews-*.csv,articles-*.csv) and their normalized tables likewise carry reviewer and author names, emails, and ORCIDs. They write underOJS_DOWNLOADS_DIR(defaultdata/ojs-website, under the gitignoreddata/) — apply the same umask/private-directory care as for the API dumps.
Changelog
Release history is in CHANGELOG.md.
Release files for ojs 0.10.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ojs-0.10.0.tar.gz | 89.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ojs-0.10.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 185.9 kB
Release files / ojs-0.10.0.tar.gz
| Download URL | ojs-0.10.0.tar.gz |
|---|---|
| Size | 89.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b7d53dc4b6e8da7d843d73f18d6fa59e6f25c10f846a5eb0f793b5c202a896a8
|
|
BLAKE2b-256 checksum How to use checksums |
60d47b17061d1d56568f1bef61d57e921df105df7b29f560ab90e7ce6de5b70c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.
Transparency logRelease files / ojs-0.10.0-py3-none-any.whl
| Download URL | ojs-0.10.0-py3-none-any.whl |
|---|---|
| Size | 96.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
5a1769838e127c2ae8053b15e9bd6882f7d9921ab2cdd69655ff3eb765eb2997
|
|
BLAKE2b-256 checksum How to use checksums |
fde7f97dc79d260dc14b1d0478231c7534643cc8e7d308ca1010228728526fda
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.
Transparency log