Dataset Explorer MCP
Understand a dataset by asking your AI assistant questions about it. This small Python server reads files on your computer and calculates the answers using Pandas. It checks missing values, finds duplicates, summarizes columns, and shows relationships between features.
It runs locally through stdio: your MCP client starts the server and talks to it directly. You don't need a website, a hosting account, or a Gemini API key. The server reads your files without changing them. Your assistant may send tool results to its AI provider according to that client's settings.
Install
You need Python 3.10 or newer.
python -m pip install dataset-explorer-mcp
The command to start the server is:
dataset-explorer-mcp
Your MCP client normally runs this command for you. If you run it in a terminal, it waits quietly for messages from a client. That is expected.
If you use uv, you can run it without a separate installation:
uvx dataset-explorer-mcp
Connect your assistant
Use an MCP client that supports local stdio servers, such as Claude Desktop, VS Code with Copilot, or Cursor. The settings file differs by client.
For Claude Desktop or Cursor, add this to your MCP configuration:
{
"mcpServers": {
"dataset-explorer": {
"command": "uvx",
"args": ["dataset-explorer-mcp"]
}
}
}
For VS Code, use .vscode/mcp.json:
{
"servers": {
"dataset-explorer": {
"type": "stdio",
"command": "uvx",
"args": ["dataset-explorer-mcp"]
}
}
}
If you installed with pip, use "command": "dataset-explorer-mcp" and "args": []
instead. An absolute path to the executable also works. Reload your client after
changing its configuration.
Use a local file
Give your assistant the full path to your dataset. For example:
Explore
C:/Users/YourName/Downloads/customers.csv. Check missing values and duplicates.
Summarize
/home/yourname/data/sales.xlsxand inspect the revenue column.
In
/Users/yourname/data/results.parquet, which features are associated with the target columnscore?
All tools take a path. A direct tool call looks like:
{"path": "C:/Users/YourName/Downloads/customers.csv"}
Use forward slashes in Windows paths, or double backslashes when writing JSON. The file must be available on the computer where the server runs.
Supported files and tools
Supported files: CSV, TSV, Excel (.xlsx, .xls), JSON, and Parquet.
Excel reads the first worksheet. JSON must contain tabular data that Pandas can read.
| Tool | What it does |
|---|---|
get_dataset_overview |
Lists columns, types, and missing-value counts |
dataset_shape |
Counts rows and columns |
dataset_statistical_summary |
Calculates numeric means and medians |
inspect_Column |
Summarizes one column; also takes col_name |
analyze_target |
Inspects a target column; also takes target_name |
duplicate_finder |
Finds repeated rows |
analyze_missing_values |
Reports missing data |
find_correlations |
Finds related numeric columns; optional threshold defaults to 0.8 |
detect_outliers |
Finds unusual numeric values |
screen_target_relationships |
Compares features with a target; also takes target |
The server also offers the dataset://guide resource and an explore_dataset
prompt. These results help you explore data; they don't prove causes or train a model.
Large files need enough RAM because each tool loads the dataset into memory.
Troubleshooting
- Command not found: use the full path to
dataset-explorer-mcp, or install uv and use theuvxconfiguration above. - File not found: use an absolute path and check that the server can read it.
- No tools appear: check your client's server logs and reload its MCP settings.
- Server seems idle: it is waiting for the MCP client; connect it through your assistant rather than typing questions into the server terminal.
- Unsupported file: save the data in one of the formats listed above.
- Missing values in statistics: empty or constant columns may have undefined statistics. Check the overview and missing-value tools first.
Normal server output is reserved for MCP messages. Diagnostics go to stderr, which your client's server logs usually display.
Run from source
git clone https://github.com/khanarmaghanrasheed-18/MCP-Dataset-Explorer.git
cd MCP-Dataset-Explorer
python -m pip install -e ".[dev]"
python -m pytest -q
python mcp_server.py
Build the downloadable package with python -m build.
License
MIT. See LICENSE.
Metadata
Release files for dataset-explorer-mcp 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| dataset_explorer_mcp-0.2.0.tar.gz | 16.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| dataset_explorer_mcp-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 27.3 kB
Release files / dataset_explorer_mcp-0.2.0.tar.gz
| Download URL | dataset_explorer_mcp-0.2.0.tar.gz |
|---|---|
| Size | 16.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
794ef7bee93418c9a2ecfb3bbd647c30e35a35fc92b96689c143de47a0c04aad
|
|
BLAKE2b-256 checksum How to use checksums |
c0530f454faba5c80f536a97c9cf9d827f789a29d30a665ad643f5e4daa3af94
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.4
|
Release files / dataset_explorer_mcp-0.2.0-py3-none-any.whl
| Download URL | dataset_explorer_mcp-0.2.0-py3-none-any.whl |
|---|---|
| Size | 10.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a3319bdc0d9db702f28f82bd583a9602a337f8f384e47a86de63b6a5f4cb2886
|
|
BLAKE2b-256 checksum How to use checksums |
c241feb83e20c6d0449c2aae2be62d67bf8a528851cb10ddeb1a30878ecfcd95
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.4
|