Skip to main content

A simple simulator of a system which implements map/reduce paradigm.

Project description

MREDU

A simple simulator of a system which implements map/reduce paradigm similarly to how Apache Hadoop does. Its objective is to be used as an educational tool to learn how to code map/reduce algorithms without needing to install complex components.

Requirements

To use it you will need:

  • A Python 3.8+ interpreter.

  • uv package manager (recommended) or pip

  • Install required packages from PyPi:

    With uv (recommended):

    uv sync
    

    Or with pip:

    pip install mredu
    

Examples

There are several examples of use of the simulator in the examples directory.

  • example1: Does some calculations from a list of tuples.
  • example2: Calculates the histogram of the number of words per line in the file quijote.txt from data folder.
  • example3: The ubiquitous word-count example written to run on the simulator and applied to the same quijote.txt file.
  • example4: Inverse k,v -> v,k agrupation

To run an example with uv:

uv run python examples/example3.py

Development

To contribute to this project, you will need to set up a development environment.

  1. Clone the repository:

    git clone https://github.com/ramonpin/mredu.git
    cd mredu
    
  2. Install uv (if not already installed):

    curl -LsSf https://astral.sh/uv/install.sh | sh
    
  3. Install dependencies:

    Install all dependencies including dev dependencies:

    uv sync
    

    This will:

    • Create a virtual environment in .venv/
    • Install the package in editable mode
    • Install all runtime and development dependencies
    • Create a uv.lock file for reproducible builds
  4. Run the tests:

    To make sure everything is working correctly, run the test suite:

    uv run pytest
    

    To run tests with coverage:

    uv run pytest --cov=mredu
    

Docs

mredu simulates a MapReduce environment. The process is as follows:

  1. Input: You start with an input sequence of (key, value) pairs. mredu provides helper functions to read data from files into this format.
  2. Map: A mapper function is applied to each (key, value) pair, producing a new sequence of (key, value) pairs.
  3. Shuffle & Sort: The framework automatically groups the pairs from the map phase by key.
  4. Reduce: A reducer function is applied to each key and its list of associated values, producing the final result.

Core Functions

  • input_file(path): Reads a text file line by line, producing a sequence of (line_number, line_content) pairs.
  • input_kv_file(path, sep): Reads a text file line by line, splitting each line by sep to produce (key, value) pairs.
  • map_red(input_sequence, mapper, reducer): Chains together the map, shuffle/sort, and reduce steps. It takes an input sequence and the mapper and reducer functions as arguments.
  • run(map_red_process): Executes the full MapReduce process and prints the resulting (key, value) pairs to the console, separated by a tab.

Example: Word Count

Here is how you would implement the classic word count example using mredu.

First, you define your mapper function. It takes a key and a value as input (in this case, line number and line text). It splits the line into words, and for each word, it returns a (word, 1) pair.

import re

def mymap(_, v):
    words = list(filter(lambda s: s != '', re.split(r'\W', v)))
    return [(word.lower(), 1) for word in words]

Next, you define your reducer function. It takes a key (a word) and a list of values (a list of 1s) and returns a pair with the word and the sum of the values.

def myred(k, vs):
    return k, len(vs)

Finally, you tie it all together. You create an input source from a file, pass it to map_red with your mapper and reducer, and then use run to execute the process.

from mredu.simul import map_red, input_file, run

# assuming mymap and myred are defined as above

if __name__ == '__main__':
    process = map_red(input_file('data/quijote.txt'), mymap, myred)
    run(process)

This will output the word counts to the console.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mredu-1.1.0.tar.gz (213.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mredu-1.1.0-py3-none-any.whl (8.9 kB view details)

Uploaded Python 3

File details

Details for the file mredu-1.1.0.tar.gz.

File metadata

  • Download URL: mredu-1.1.0.tar.gz
  • Upload date:
  • Size: 213.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.1

File hashes

Hashes for mredu-1.1.0.tar.gz
Algorithm Hash digest
SHA256 98ecfdb4fe47d1baa743a8c0d248d05d9b0b1a3f8e9a72c20381b65e01462b38
MD5 7410fc0dba8d9422a8b531949b0d1ed0
BLAKE2b-256 b4e94fe4b88a1867e553c3a583eb4d74ee49689f0ceeee9da3cc2c96d1965e2f

See more details on using hashes here.

File details

Details for the file mredu-1.1.0-py3-none-any.whl.

File metadata

  • Download URL: mredu-1.1.0-py3-none-any.whl
  • Upload date:
  • Size: 8.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.1

File hashes

Hashes for mredu-1.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 58d2479fcbac300579922ffa19d7190f3df0015070028ec268c951fbebb78fc1
MD5 cd1f3c5b1e3feefb501a7f5c08cfe638
BLAKE2b-256 1a7164fd9c2dae6dbf8e8439289f31d0b45fd6c93d1e849752f6e3747ca8d524

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page