Query JSON files using Elasticsearch Query DSL syntax
Project description
JSONSHQUERY
Query JSON files using Elasticsearch Query DSL syntax.
Overview
Jsonshquery lets you query array of object in a json file or python dictionary by following Elasticsearch query DSL. If you are familiar with Elasticsearch, then there shouldn't be any problem for you to use this tool.
Features
- Query JSON files (arrays of objects) using Elasticsearch Query DSL
- Support for basic query clauses:
term,terms,match,match_phrase,multi_match,prefix,exists - Boolean queries:
must,must_not,should,filter - Nested field support using dot notation (e.g.,
address.city) - Interactive CLI with JSON syntax highlighting and auto-completion
Installation
pip install jsonshquery
Usage
Interactive CLI
To start interacting with the command line, run this command.
jsonshquery --file <./path/to/file.json>
Look at this example. Given a JSON file data.json:
[
{
"id": 2,
"name": "Product B",
"category": "Furniture",
"price": 89.50,
"in_stock": true,
"tags": ["wood"],
"supplier": {"name": "Supplier Y", "country": "Indonesia"},
"description": "A durable wooden furniture piece crafted with quality materials, suitable for home interiors, providing comfort, stability, and aesthetic appeal for daily use."
},
{
"id": 3,
"name": "Product C",
"category": "Clothing",
"price": 29.99,
"in_stock": true,
"tags": ["fashion"],
"supplier": {"name": "Supplier Z", "country": "Vietnam"},
"description": "A stylish clothing item made with breathable fabric, designed for comfort and versatility, suitable for casual wear and adaptable to various weather conditions."
},
{
"id": 5,
"name": "Product E",
"category": "Furniture",
"price": 59.99,
"in_stock": true,
"tags": ["sale"],
"supplier": {"name": "Supplier A", "country": "Indonesia"},
"description": "A compact furniture unit designed for comfort, offering functional storage, lightweight structure, and modern appearance suitable for apartments and offices."
}
]
Start the interactive CLI:
jsonshquery --file data.json
This Elasticsearch-style query demonstrates the capabilities of jsonshquery:
{
"result_path" : "./result/query_test.json",
"_source" : ["id", "name", "category", "description"],
"query": {
"bool": {
"must": [
{ "match" : { "description": "comfort" }},
{ "range" : { "price" : { "gte" : 40, "lte" : 100 }}},
{ "term" : { "in_stock" : true }},
{ "term" : { "supplier.country" : "Indonesia" }}
],
"must_not": [
{ "term": { "category": "Clothing" }}
]
}
}
};
End your query with ; and press Enter to execute. When you use this interactive CLI the result of the query will be created in a json file.
The result_path field is optional, you use it to specify in which directory the query result will be created, if you don't specify it by default the file name will be jsonshquery_result.json. This is the expected output format:
{
"count": 2,
"hits": [
{
"id": 2,
"name": "Product B",
"category": "Furniture",
"description": "A durable wooden furniture piece crafted with quality materials, suitable for home interiors, providing comfort, stability, and aesthetic appeal for daily use."
},
{
"id": 5,
"name": "Product E",
"category": "Furniture",
"description": "A compact furniture unit designed for comfort, offering functional storage, lightweight structure, and modern appearance suitable for apartments and offices."
}
]
}
Python Module
You can also use jsonshquery as module
from jsonshquery import Jsonshquery
data = [
{"id": 1, "name": "Product A", "category": "Furniture", "price": 99.99},
{"id": 2, "name": "Product B", "category": "Electronics", "price": 49.99}
]
query = {
"query": {
"bool": {
"must": [
{"term": {"category": "Furniture"}}
]
}
}
}
jq = Jsonshquery(data)
result = jq.search_by_query(query)
print(f"Found {result['count']} products")
Supported Query Types
Leaf Query Clauses
- term: Exact field match
- terms: Field matches one of multiple values
- match: Full-text search
- match_phrase: Phrase matching
- multi_match: Search across multiple fields
- match_all: just return all documents
- prefix: Prefix matching
- match_phrase_prefix : Combination of
match_phraseandprefix - range : Field value within range
- exists: Check if field exists
- wildcard : Match by pattern (*, ?, etc)
- regexp : Match by regex
Boolean Queries
- must: All clauses must match (AND)
- must_not: All clauses must not match (NOT)
- should: At least one clause must match (OR)
- filter: Same as must, but no scoring
Other Keywords
- _source: Fields that wanted to be returned, if not specified it will return complete documents.
- size: Limit the number of matching documents to return, if not specified it will return all matching documents.
Some Important Notes
Native elasticsearch implements scoring on query result, on the other hand, jsonshquery doesn't do that. Every matched documents are purely based on the query given without any scoring process, so it is more like filtering. Therefore keyword filter and must actually works the same here.
In Elasticsearch, if should keyword is being used together with must it will become a scoring booster to the documents that matched with must clause, but since jsonshquery doesn't care about scoring, using should with must is kinda useless and the codebase itself ignore that. Anyway if should is the only boolean clause used in bool, jsonshquery and elasticsearch will work pretty equal, that is, it becomes an OR statement for the leaf clauses inside should.
range clause currently limited to only value with numerical type, you cannot do date or time filtering using range here. nested keyword is also not available.
jsonshquery is not yet optimized for speed, our current objective is to make it works as expected. But you generally won't notice long query execution time for even medium to high size documents.
This tool will load your json data into memory, so you might have some consideration when your json file has a huge size.
Possible Errors
- Data is not in form list of objects/dictionaries
- Invalid JSON body
- Malformed query body
License
MIT License - see LICENSE file for details.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file jsonshquery-1.0.3.tar.gz.
File metadata
- Download URL: jsonshquery-1.0.3.tar.gz
- Upload date:
- Size: 15.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7734b39d1b7341fa7ce9b22bbcf3c44ef9962a72a7e6e7a1348bbaaca1a67011
|
|
| MD5 |
ced8060c7d5709f7d61faf6a5b188872
|
|
| BLAKE2b-256 |
cb75beb651cc5856f58153f05ea8fb7c59541b9650e98fd3c04cb55c2c9722f1
|
File details
Details for the file jsonshquery-1.0.3-py3-none-any.whl.
File metadata
- Download URL: jsonshquery-1.0.3-py3-none-any.whl
- Upload date:
- Size: 11.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.10.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ec64a90a2c0ea7d781e2898be79d84cdcde62f99d97088552f4470a9996671f4
|
|
| MD5 |
53bb59ea6f959bcde2ac9451ee786f46
|
|
| BLAKE2b-256 |
ced0f794b408180f3b7585073c382fb07e5dccdd1dd4dad01f614390d90661f9
|