Agent-friendly Spark Connect CLI: read-only querying + async long-job control. No JVM, no Kerberos on the client.

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

runningOtter

These details have not been verified by PyPI

Project description

spark-connect-cli (`scq`)

An agent-friendly Spark Connect CLI — read-only querying plus async control for long-running jobs.

Built for LLM agents and humans who live in a shell. Unlike spark-sql / spark-submit, the client is a thin pure-Python gRPC client: no JVM, and no Kerberos on the client side — the Spark Connect server authenticates with its own keytab, so you just point at sc://host:15002 and go.

Why

JSON-first, read-only by default. Safe for an agent to call for exploration; writes/DDL are blocked unless you opt in (--allow-ddl).
Long jobs don't block you. A multi-minute Spark job shouldn't trap an agent in a 30-minute tool call. scq submits the job, hands back a durable job id, and returns immediately. Poll it whenever you like; the handle survives a client/container restart because it lives in an on-disk registry.
Stable exit codes so a caller can branch without scraping text.

Install

pip install spark-connect-cli         # once published
# or, from source:
pip install -e .

Quick start

export SPARK_REMOTE=sc://localhost:15002   # your Spark Connect endpoint

scq databases
scq tables mydb --like '%orders%'
scq describe mydb.orders
scq query "SELECT id, name FROM mydb.orders LIMIT 10"

Output is JSONEachRow (one JSON object per line) by default; pick another with --format json|csv|tsv|table.

Read-only guard

scq query allows only SELECT/SHOW/DESCRIBE/EXPLAIN/WITH. Anything else exits with code 3 unless you pass --allow-ddl.

exit	meaning
0	success
1	query error (bad SQL)
2	connection error
3	blocked by the read-only guard
4	job-control error (no such job, …)

Async jobs (Layer A)

Long work runs detached and is tracked by a file-based registry under $SCQ_JOBS_DIR (default ~/.spark-connect-cli/jobs).

# submit — returns a job id immediately, does NOT block
scq sync ods.orders --to clickhouse
# {"job_id": "j-20260625-...", "state": "running", "message": "... poll with ..."}

scq jobs list                       # all jobs + state
scq jobs status j-20260625-...      # full status (rows, timings, pid, exit code)
scq jobs logs   j-20260625-... --tail 40
scq jobs cancel j-20260625-...      # kills the whole process group

Design: each job is a directory with meta.json (state machine: submitted → running → succeeded|failed|cancelled) and out.log. The worker runs in its own process group, so cancel kills the entire tree (no orphans). A running job whose process has vanished is reconciled to failed on the next status read, so status never lies.

Hive → ClickHouse sync

scq sync is one job kind built on the async subsystem. It uses Spark direct write: a Spark Connect job reads the Hive table and writes to ClickHouse over JDBC. The write runs on the executors, so rows never pass through this process or the agent.

Modes control write parallelism — single (one connection, small tables), parallel (N partitions, large tables), auto (picks by row count).

Requires:

clickhouse-jdbc on the Spark Connect server classpath (/opt/spark/jars/),
cluster→ClickHouse network egress,
a JDBC URL with credentials via --ch-jdbc / $SCQ_CH_JDBC,
the target ClickHouse table created beforehand with a suitable engine (Spark append won't build a usable MergeTree table for you — create it first, e.g. with the chsql skill).

Introspection

scq meta db.table            # one JSON: schema, created time, location,
                             # partitions, file count/size, mtime range
scq meta db.table --count    # also run an exact count(*)

scq exec stages?status=active            # read-only Spark REST passthrough
scq exec executors
scq exec stages/<id>/<attempt>/taskSummary?quantiles=0.5,0.95,1.0   # skew: max/median

scq exec auto-discovers the running Spark app via the YARN ResourceManager and proxies its monitoring REST API (GET-only). Set the RM base with $SCQ_YARN_RM.

Reading scq exec executors — the maxMemory field is Spark's storage/cache pool ((heap − 300 MB reserved) × 0.6), not the executor's total memory: a 512 MB executor reports ~93 MB, a 1536 MB driver ~741 MB. The real heap is spark.executor.memory (+ off-heap overhead). The driver row has 0 cores and runs no tasks. With dynamic allocation, idle executors are released — so the list may show only the driver when nothing is running.

Configuration

env	default	meaning
`SPARK_REMOTE`	`sc://localhost:15002`	Spark Connect endpoint
`SCQ_JOBS_DIR`	`~/.spark-connect-cli/jobs`	job registry (put on a persistent volume)
`SCQ_MAX_ROWS`	`10000`	default row cap for `query`
`SCQ_CH_JDBC`	—	ClickHouse JDBC URL for `sync` path A
`SCQ_YARN_RM`	`http://namenode.hive-net:8088`	YARN RM base for `scq exec`

Use with an LLM agent

SKILL.md ships a ready-made skill (discover-before-query workflow, async-job etiquette, type-mapping table). Drop it into your agent's skills directory and the agent drives scq through a shell/Bash tool.

Roadmap

Clarify in SKILL.md that scq exec executors maxMemory is the storage pool, not total memory (already noted above).
scq cluster — optional read-only passthrough to the YARN ResourceManager REST (apps / queues / nodes), rounding out the introspection plane.
Vendored/offline install path (bundle wheels) for air-gapped deployments.

License

MIT

Project details

These details have been verified by PyPI

Project links

GitHub Statistics

Maintainers

runningOtter

These details have not been verified by PyPI

Release history Release notifications | RSS feed

0.3.1

Jul 3, 2026

0.3.0

Jun 30, 2026

0.2.2

Jun 26, 2026

This version

0.2.1

Jun 26, 2026

0.2.0

Jun 26, 2026

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

spark_connect_cli-0.2.1.tar.gz (18.9 kB view details)

Uploaded Jun 26, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

spark_connect_cli-0.2.1-py3-none-any.whl (22.6 kB view details)

Uploaded Jun 26, 2026 Python 3

File details

Details for the file spark_connect_cli-0.2.1.tar.gz.

File metadata

Download URL: spark_connect_cli-0.2.1.tar.gz
Upload date: Jun 26, 2026
Size: 18.9 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for spark_connect_cli-0.2.1.tar.gz
Algorithm	Hash digest
SHA256	`3210ca754b1144db74be90680d09e090cf930448479dd6501749c2d4c30e11c8`
MD5	`f196ae02f116cc2b7e32366506118cf7`
BLAKE2b-256	`f6b54fda457c2fb31d4798c80b510a95beaa194b16f0752cecb406c0f88a8188`

See more details on using hashes here.

Provenance

The following attestation bundles were made for spark_connect_cli-0.2.1.tar.gz:

Publisher: publish.yml on dengshu2/spark-connect-cli

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: spark_connect_cli-0.2.1.tar.gz
- Subject digest: 3210ca754b1144db74be90680d09e090cf930448479dd6501749c2d4c30e11c8
- Sigstore transparency entry: 1963491851
- Sigstore integration time: Jun 26, 2026
Source repository:
- Permalink: dengshu2/spark-connect-cli@4d6732a8148b29b650b7c0a4199b1768890400a8
- Branch / Tag: refs/tags/v0.2.1
- Owner: https://github.com/dengshu2
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@4d6732a8148b29b650b7c0a4199b1768890400a8
- Trigger Event: release

File details

Details for the file spark_connect_cli-0.2.1-py3-none-any.whl.

File metadata

Download URL: spark_connect_cli-0.2.1-py3-none-any.whl
Upload date: Jun 26, 2026
Size: 22.6 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for spark_connect_cli-0.2.1-py3-none-any.whl
Algorithm	Hash digest
SHA256	`7a44ba6c27b099d47f0084937a803c3b877d1de6834d3950ba4290fb5f53ce16`
MD5	`bdcd3a5e31bba3d7184bce93cb73f79b`
BLAKE2b-256	`70ef79ea7577dce218f632df38c72c42ee35928db47bf4ce62253735294edad2`

See more details on using hashes here.

Provenance

The following attestation bundles were made for spark_connect_cli-0.2.1-py3-none-any.whl:

Publisher: publish.yml on dengshu2/spark-connect-cli

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: spark_connect_cli-0.2.1-py3-none-any.whl
- Subject digest: 7a44ba6c27b099d47f0084937a803c3b877d1de6834d3950ba4290fb5f53ce16
- Sigstore transparency entry: 1963491980
- Sigstore integration time: Jun 26, 2026
Source repository:
- Permalink: dengshu2/spark-connect-cli@4d6732a8148b29b650b7c0a4199b1768890400a8
- Branch / Tag: refs/tags/v0.2.1
- Owner: https://github.com/dengshu2
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@4d6732a8148b29b650b7c0a4199b1768890400a8
- Trigger Event: release

spark-connect-cli 0.2.1

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Classifiers

Project description

spark-connect-cli (scq)

Why

Install

Quick start

Read-only guard

Async jobs (Layer A)

Hive → ClickHouse sync

Introspection

Configuration

Use with an LLM agent

Roadmap

License

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance

spark-connect-cli (`scq`)