Skip to main content

LeetScrape

Python application deploy-docs PYPI

Introducing the LeetScrape - a powerful and efficient Python package designed to scrape problem statements and basic test cases from LeetCode.com. With this package, you can easily download and save LeetCode problems to your local machine, making it convenient for offline practice and studying. It is perfect for software engineers and students preparing for coding interviews. The package is lightweight, easy to use and can be integrated with other tools and IDEs. With the LeetScrape, you can boost your coding skills and improve your chances of landing your dream job.

Use this package to get the list of Leetcode questions, their topic and company tags, difficulty, question body (including test cases, constraints, hints), and code stubs in any of the available programming languages.

Detailed documentation available here.

Installation

Start by installing the package from pip or conda:

pip install leetscrape
# or using conda:
conda install leetscrape
# or using poetry:
poetry add leetscrape

Usage

Command Line

After installing the package, run the following command to get a code stub and a pytest test file for a given Leetcode question:

$ leetscrape --titleSlug two-sum --qid 1

At least one of the two arguments is required.

  • titleSlug is the slug of the leetcode question that is in the url of the question, and
  • qid is the number associated with the question.

Other classes

Import the relevant classes from the package:

from leetscrape.GetQuestionsList import GetQuestionsList
from leetscrape.GetQuestionInfo import GetQuestionInfo
from leetscrape.utils import combine_list_and_info, get_all_questions_body

Scrape the list of problems

Get the list of questions, companies, topic tags, categories using the GetQuestionsList class:

ls = GetQuestionsList()
ls.scrape() # Scrape the list of questions
ls.to_csv(directory_path="../data/") # Save the scraped tables to a directory

Get Question statement and other information

Query individual question's information such as the body, test cases, constraints, hints, code stubs, and company tags using the GetQuestionInfo class:

# This table can be generated using the previous commnd
questions_info = pd.read_csv("../data/questions.csv")

# Scrape question body
questions_body_list = get_all_questions_body(
    questions_info["titleSlug"].tolist(),
    questions_info["paidOnly"].tolist(),
    save_to="../data/questionBody.pickle",
)

# Save to a pandas dataframe
questions_body = pd.DataFrame(
    questions_body_list
).drop(columns=["titleSlug"])
questions_body["QID"] = questions_body["QID"].astype(int)

Note The above code stub is time consuming (10+ minutes) since there are 2500+ questions.

Create a new dataframe with all the questions and their metadata and body information.

questions = combine_list_and_info(
    info_df = questions_body, list_df=ls.questions, save_to="../data/all.json"
)

Upload scraped data to a Database

Create a PostgreSQL database using the SQL dump and insert data using sqlalchemy.

from sqlalchemy import create_engine
from sqlalchemy.orm import sessionmaker

engine = create_engine("<database_connection_string>", echo=True)
questions.to_sql(con=engine, name="questions", if_exists="append", index=False)
# Repeat the same for tables ls.topicTags, ls.categories,
# ls.companies, # ls.questionTopics, and ls.questionCategory

Use the queried_questions_list PostgreSQL function (defined in the SQL dump) to query for questions containy query terms:

select * from queried_questions_list('<query term>');

Use the all_questions_list PostgreSQL function (defined in the SQL dump) to query for all the questions in the database:

select * from all_questions_list();

Use the get_similar_questions PostgreSQL function (defined in the SQL dump) to query for all questions similar to a given question:

select * from get_similar_questions(<QuestionID>);

Use the extract_solutions method to extract solution code stubs from your python script. Note that the solution method should be a part of a class named Solution (see here for an example):

# Returns a dict of the form {QuestionID: solutions}
solutions = extract_solutions(filename=<path_to_python_script>)

Use the upload_solutions method to upload the extracted solution code stubs from your python script to the PosgreSQL database.

upload_solutions(engine=<sqlalchemy_engine>, row_id = <row_id_in_table>, solutions: <solutions_dict>)

Metadata

Release files for leetscrape 0.1.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for leetscrape 0.1.5
File Size Uploaded
leetscrape-0.1.5.tar.gz 14.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for leetscrape 0.1.5
File Interpreter ABI Platform
leetscrape-0.1.5-py3-none-any.whl Python 3 none any Details

Total release size: 29.0 kB

Release files / leetscrape-0.1.5.tar.gz

Download URL leetscrape-0.1.5.tar.gz
Size 14.6 kB
Tags Source
SHA-256 checksum
How to use checksums
268ca726e84da524cbdef553d7553232e20f6400dbb510ab79194f1026e9434f
BLAKE2b-256 checksum
How to use checksums
76a288038a4c60ea2c4731bf80922ec15c5b7bd5922256769d653b80a239d6ca
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.3.2 CPython/3.10.7 Windows/10

Release files / leetscrape-0.1.5-py3-none-any.whl

Download URL leetscrape-0.1.5-py3-none-any.whl
Size 14.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3aade3a914a25888169c11e2fc4e0bd0f190a1b56c5278922ef78d6a4e7eafca
BLAKE2b-256 checksum
How to use checksums
9c877168cbbc8fb310e2b44fd5f8f13bf670cacab38513aaa8e062e64c97504e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via poetry/1.3.2 CPython/3.10.7 Windows/10

Release history Release notifications | RSS feed

1.0.1

2 release files

1.0.0

2 release files

0.1.11

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.6

2 release files

This release

0.1.5 This release

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page