Skip to main content

Faster and more Insightful analysis of survey results

This package lets you apply advanced Natural Language Processing (NLP) and Machine Learning functions on survey results directly within a dataframe.

It fills a gap where many NLP packages (like spacy, genism, sentence_transformers) are not designed for data in a spreadsheet (and therefore imported into a dataframe), and many of the people who are tasked with analysing survey results are often not data scientists.

For example, to extract the sentiment you can just type:

df.extract_sentiment(input_column="survey-comments")

It will abstract away a lot of the data transformation pipeline to give you useful functionality with minimal code.

Examples

See ReadTheDocs for simple example notebooks. There are more detailed notebooks in the repo under notebooks/

Functionality

Clustering comments

It will group similar free-text comments together and assign a cluster ID. This is a useful step prior to any qualitative analysis.

Sentiment Analysis

It will measure the sentiment in terms or postive / neutral / negative and assign a score for each of those parts, picking the highest scoring as the most likely overall sentiment.

Topic analysis

Involves TFIDF and word co-occurence to gain some high level insights into the likely topics

Clustering respondents by their Likert answers

For strongly disagree ... neutral ... strongly agree type responses, this groups respondents who answer along similar lines, which can be far more useful than overall averages across the survey.

Two approaches are available:

  • df.cluster_respondents(...) — UMAP (cosine) + HDBSCAN. Best when you have many respondents.
  • df.cluster_respondents_correlation(...) — respondent correlation + hierarchical (dendrogram) clustering. Best when you have few respondents relative to the number of questions, because each correlation is estimated across all the questions.

encode_likert(..., scale=5) keeps the intensity of agreement (strongly agree/disagree → ±2). See docs/source/clustering_methods_comparison.md for the maths behind the two methods and guidance on encoding and small surveys. (cluster_questions is a deprecated alias for cluster_respondents.)

Visualisation

Functions to help make sense of the clusters and topics you have identified using the above functions (in development)

Setup

If sentence transformers throws dll errors: https://stackoverflow.com/questions/78484297/c-torch-lib-fbgemm-dll-or-one-of-its-dependencies/78794748#78794748

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pandas_survey_toolkit-1.3.0.tar.gz (611.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pandas_survey_toolkit-1.3.0-py3-none-any.whl (26.0 kB view details)

Uploaded Python 3

File details

Details for the file pandas_survey_toolkit-1.3.0.tar.gz.

File metadata

  • Download URL: pandas_survey_toolkit-1.3.0.tar.gz
  • Upload date:
  • Size: 611.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pandas_survey_toolkit-1.3.0.tar.gz
Algorithm Hash digest
SHA256 b200c4e1c23e9ca3d0f14bf3f87ce509f35a45d527a27e309f8b310258383ebf
MD5 a450c6ecfb3e19984d6e8f270017934b
BLAKE2b-256 470ad91e0d648a4484776933cd00141c668a04124f221a58b46f9feb1d03d555

See more details on using hashes here.

File details

Details for the file pandas_survey_toolkit-1.3.0-py3-none-any.whl.

File metadata

File hashes

Hashes for pandas_survey_toolkit-1.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 056a2f154c6235485148564a0870a3a5587d34faaf1e3a02e4257b88ab313fd4
MD5 8cb092a49441908c96e620f49453cc5d
BLAKE2b-256 88d20959d84dc86859c73f8411ced485b38a73902813ef9f547fd27c0fc3efa2

See more details on using hashes here.

Release history Release notifications | RSS feed

2.0.1

2 files

This release

1.3.0 This release

2 files

1.2.0

2 files

1.1.1

2 files

1.1

2 files

1.0.14

2 files

1.0.13

2 files

1.0.12

2 files

1.0.11

2 files

1.0.10

2 files

1.0.9

2 files

1.0.4

2 files

1.0.3

2 files

1.0.2

2 files

1.0.1

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page