Skip to main content

Faster and more Insightful analysis of survey results

This package lets you apply advanced Natural Language Processing (NLP) and Machine Learning functions on survey results directly within a dataframe.

It fills a gap where many NLP packages (like spacy, genism, sentence_transformers) are not designed for data in a spreadsheet (and therefore imported into a dataframe), and many of the people who are tasked with analysing survey results are often not data scientists.

For example, to extract the sentiment you can just type:

df.extract_sentiment(input_column="survey-comments")

It will abstract away a lot of the data transformation pipeline to give you useful functionality with minimal code.

Installation

pip install pandas-survey-toolkit           # Likert clustering + text preprocessing (torch-free)
pip install "pandas-survey-toolkit[nlp]"    # + free-text comment tooling (sentence embeddings, spaCy, transformer sentiment)

The core install is lightweight and does not pull in torch/spacy, so clustering Likert questions/respondents works without the heavy deep-learning stack. Install the [nlp] extra when you need free-text comment embeddings, spaCy pipelines, or transformer sentiment.

Examples

See ReadTheDocs for simple example notebooks. There are more detailed notebooks in the repo under notebooks/

Functionality

Clustering comments

It will group similar free-text comments together and assign a cluster ID. This is a useful step prior to any qualitative analysis.

Sentiment Analysis

It will measure the sentiment in terms or postive / neutral / negative and assign a score for each of those parts, picking the highest scoring as the most likely overall sentiment.

Topic analysis

Involves TFIDF and word co-occurence to gain some high level insights into the likely topics

Clustering a Likert survey (respondents and questions)

For strongly disagree ... neutral ... strongly agree responses, the toolkit clusters both axes purely from the data (no NLP on the question text):

Responses are encoded on a simple 3-point scale (−1 / 0 / +1) and grouped with cosine distance, which compares the direction of people's opinions — so respondents who agree with everything and respondents who disagree with everything separate cleanly.

  • Respondents — group people who answer along similar lines:
    • df.cluster_respondents(method="auto") — picks the engine by survey size (below). This is the one to reach for.
    • df.cluster_respondents_cosine(...) — cosine distance + hierarchical (dendrogram) clustering. Best for small/medium surveys.
    • df.cluster_respondents_umap(...) — UMAP + HDBSCAN. Scales to very large surveys (many respondents).
  • Questions — group questions that get answered similarly:
    • df.cluster_questions(...) — returns a pd.Series (question → cluster id) you can inspect or save to CSV. (In 2.0 this clusters questions; in 1.x it was a deprecated alias that clustered respondents.)
  • Everything at oncedf.cluster_survey(likert_cols) clusters respondents and questions and orders both axes so the plots below "just work".

See docs/source/clustering_methods_comparison.md for the maths, the distance-threshold rule of thumb, and small-survey advice.

Visualisation

  • cluster_heatmap_plot(...) — Altair heatmap of sentiment per respondent-cluster × question, with a cluster-size bar chart. Best for many respondents / few clusters; questions auto-ordered by their cluster.
  • survey_clustermap(...) — seaborn biclustered clustermap of individual respondents × questions with marginal dendrograms. Best for few respondents, where you can see everyone.
  • plot_respondent_dendrogram(...) — the respondent dendrogram from cosine clustering.

Setup

If sentence transformers throws dll errors: https://stackoverflow.com/questions/78484297/c-torch-lib-fbgemm-dll-or-one-of-its-dependencies/78794748#78794748

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pandas_survey_toolkit-2.0.1.tar.gz (620.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pandas_survey_toolkit-2.0.1-py3-none-any.whl (33.3 kB view details)

Uploaded Python 3

File details

Details for the file pandas_survey_toolkit-2.0.1.tar.gz.

File metadata

  • Download URL: pandas_survey_toolkit-2.0.1.tar.gz
  • Upload date:
  • Size: 620.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pandas_survey_toolkit-2.0.1.tar.gz
Algorithm Hash digest
SHA256 556cd5ffdb7e0e51b41d1347453f195492e9b5941b678c599b64302c7cccb464
MD5 0472de01cd503f5a53cce9f12d54c239
BLAKE2b-256 12b4416f1866e317681fd1d10c2a7a95985f3fe5fb2b040ad5b96df0f7ea5086

See more details on using hashes here.

File details

Details for the file pandas_survey_toolkit-2.0.1-py3-none-any.whl.

File metadata

File hashes

Hashes for pandas_survey_toolkit-2.0.1-py3-none-any.whl
Algorithm Hash digest
SHA256 d1ee11abf60421969e65d463b3c98829ea837e89175d82b5c09698e5470abd0c
MD5 523f40581fd0cfdd152e93c77d35aaa3
BLAKE2b-256 2b26c0ceb092f893c1d581b0d996ecb98d5690f5b500fd2ab2e2cdc62282d849

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

2.0.1 This release

2 files

1.3.0

2 files

1.2.0

2 files

1.1.1

2 files

1.1

2 files

1.0.14

2 files

1.0.13

2 files

1.0.12

2 files

1.0.11

2 files

1.0.10

2 files

1.0.9

2 files

1.0.4

2 files

1.0.3

2 files

1.0.2

2 files

1.0.1

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page