Skip to main content
Archived

This project has been archived by its maintainers, and is no longer receiving any updates.

Persian Sentiment Analyzer

Python Version License

A Python library for sentiment analysis of Persian (Farsi) text, capable of classifying opinions as "recommended", "not_recommended", or "no_idea".

Features

  • Text Preprocessing: Normalization, tokenization, stemming, and stopword removal for Persian text
  • Word Embeddings: Built-in Word2Vec implementation for Persian language
  • Sentiment Classification: Logistic Regression classifier trained on Persian sentiment data
  • Model Persistence: Save and load trained models for future use
  • Batch Processing: Analyze sentiment for multiple texts at once

Installation

pip install persian-sentiment-analyzer

Dependencies

  • Python 3.6+

  • hazm

  • gensim

  • scikit-learn

  • numpy

  • pandas

Usage

Basic Usage

from persian_sentiment_analyzer import SentimentAnalyzer

# Initialize with a pre-trained model
analyzer = SentimentAnalyzer(model_path="path/to/pretrained_model")

# Predict sentiment
result = analyzer.predict("این محصول بسیار عالی است")
print(result)  # Output: 'recommended'

Training Your Own Model

from persian_sentiment_analyzer import SentimentAnalyzer
import pandas as pd

# Load your dataset
data = pd.read_csv("persian_reviews.csv")
texts = data['text'].tolist()
labels = data['label'].values  # 0: not_recommended, 1: recommended, 2: no_idea

# Initialize analyzer
analyzer = SentimentAnalyzer()

# Preprocess and tokenize texts
tokenized_texts = [analyzer.preprocessor.preprocess_text(text) for text in texts]

# Train Word2Vec model
analyzer.train_word2vec(tokenized_texts, vector_size=100)

# Prepare feature vectors
X = np.array([analyzer.sentence_vector(tokens) for tokens in tokenized_texts])

# Train classifier
analyzer.train_classifier(X, labels)

# Save the trained model
analyzer.save_model("my_persian_model")

Batch Processing

from persian_sentiment_analyzer import predict_sentiments_for_file

# Process a CSV file containing Persian comments
results_summary = predict_sentiments_for_file(
    analyzer,
    input_file="comments.csv",
    output_file="results.csv",
    summary_file="summary.csv"
)

print(results_summary)

Model Architecture

1- Text Preprocessing:

  • Normalization (Hazm)

  • Tokenization

  • Stemming

  • Stopword removal

2- Feature Extraction:

  • Word2Vec embeddings (100 dimensions)

  • Sentence vectors (average of word vectors)

3- Classification:

  • Logistic Regression with L2 regularization

Performance

The pre-trained model achieves the following performance on our test set:

Metric Value Accuracy 85.2% Precision 84.7% Recall 85.0% F1-score 84.8%

License This project is licensed under the MIT License - see the LICENSE file for details

Github : RezaGooner

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

persian_sentiment_analyzer-0.1.9.tar.gz (5.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

persian_sentiment_analyzer-0.1.9-py3-none-any.whl (6.3 kB view details)

Uploaded Python 3

File details

Details for the file persian_sentiment_analyzer-0.1.9.tar.gz.

File metadata

File hashes

Hashes for persian_sentiment_analyzer-0.1.9.tar.gz
Algorithm Hash digest
SHA256 0ec43de9ba14ada7180853ed1939881ee257087751f6596ae219e335e470020a
MD5 a3b66cb7e24801b2355d251737a9b7f2
BLAKE2b-256 1454afebdf16aa87a0839821aaba793adb4ce832d9adc74d46436c2637f5d322

See more details on using hashes here.

File details

Details for the file persian_sentiment_analyzer-0.1.9-py3-none-any.whl.

File metadata

File hashes

Hashes for persian_sentiment_analyzer-0.1.9-py3-none-any.whl
Algorithm Hash digest
SHA256 e21fe16465ce8c0ec705c83301fa0cb607642722f446cc2d71062dc3f3ca315e
MD5 4e84567aec2de2771dee24d27dc9e276
BLAKE2b-256 5f0a21dd7972bbb74f81c2fddb949ff36c9bdc2958da900cb21a19d6149687e6

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.1.9 This release

2 files

0.1.8

2 files

0.1.7

2 files

0.1.6

2 files

0.1.5

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page