Skip to main content

Automatic grading framework for student code in Google Colab and Jupyter notebooks

Project description

notebook-cell-tester

PyPI version Python License: MIT

Automatic grading framework for student code in Google Colab and Jupyter notebooks.

Instructors drop a single test cell below the student's code cell. The framework reads the previous cell's source from IPython history, executes it, runs the configured tests, and renders a color-coded HTML results table — no extra setup required from students.


Table of contents


Installation

pip install notebook-cell-tester

Colab one-liner — paste this at the top of the test cell if the package is not pre-installed in the environment:

!pip install -q notebook-cell-tester

Quick start

Place this cell immediately after the student's code cell in the notebook:

from notebook_cell_tester import ColabTestFramework, TestCase

tester = ColabTestFramework()

tests = [
    TestCase(
        name="add_numbers returns correct sum",
        test_type="return",
        function_name="add_numbers",
        inputs=[3, 4],
        expected=7,
        description="add_numbers(3, 4) should return 7",
    ),
    TestCase(
        name="Prints greeting",
        test_type="output",
        stdin_input="Alice",
        expected="Hello, Alice!",
    ),
]

tester.run_tests(tests)
tester.display_results()

Running this cell produces a color-coded HTML table showing each test's pass/fail status, expected vs. actual values, and (when something crashes) a collapsible traceback that students can expand.


How it works

  1. ColabTestFramework() is instantiated in the test cell — the cell below the student's code.
  2. run_tests(tests) calls load_last_cell(), which reads In[-2] from the IPython namespace to get the second-to-last cell (the student's code, not the test cell).
  3. Each TestCase is dispatched to the appropriate internal method based on test_type.
  4. display_results() renders an HTML table with pass/fail badges, result details, and collapsible error tracebacks.

Tests that only inspect the IPython namespace (variable, type_check on a variable) run even when the student hasn't executed their cell yet.


Test types

Summary

test_type Checks Key fields
return Function return value function_name, inputs, expected, tolerance
output Exact printed output stdin_input, expected, optional function_name
exception Function raises expected exception function_name, inputs, expected
regex Source code matches pattern pattern
not_regex Source code does NOT match pattern pattern
variable Variable satisfies validator variable_name, validator
type_check Value is correct type function_name or variable_name, expected
partial_output Output similarity ≥ threshold (Levenshtein) expected, similarity_threshold
regex_output Printed output matches regex pattern
contains_output Printed output contains substring expected
multiline_output Every expected line appears in output expected

return — function return value

Calls function_name(*inputs) and checks that the return value equals expected.

TestCase(
    name="Square of 5",
    test_type="return",
    function_name="square",
    inputs=[5],
    expected=25,
)

Use tolerance for floating-point comparisons (abs(result - expected) <= tolerance):

TestCase(
    name="Average of [1, 2, 3]",
    test_type="return",
    function_name="average",
    inputs=[[1, 2, 3]],
    expected=2.0,
    tolerance=1e-9,
)

output — exact stdout match

Captures everything printed to stdout and compares it to expected (after stripping leading/trailing whitespace).

# Test the whole cell — no function_name needed
TestCase(
    name="Greeting output",
    test_type="output",
    stdin_input="Alice",
    expected="Hello, Alice!",
)

# Test a specific function's printed output
TestCase(
    name="print_square(4)",
    test_type="output",
    function_name="print_square",
    inputs=[4],
    expected="16",
)

Supply stdin_input to simulate one or more input() calls. Separate multiple lines with \n.


exception — expected exception

Passes when function_name(*inputs) raises exactly the expected exception type. Fails if it returns normally or raises a different exception.

TestCase(
    name="Division by zero raises ValueError",
    test_type="exception",
    function_name="safe_divide",
    inputs=[10, 0],
    expected=ValueError,
)

regex / not_regex — source code pattern

Searches the student's source code for a Python regex. Useful for requiring or forbidding specific language constructs without running the code.

# Require a for loop
TestCase(
    name="Uses a for loop",
    test_type="regex",
    pattern=r"\bfor\b",
    description="Your solution must use a for loop.",
)

# Forbid the built-in sum()
TestCase(
    name="Does not call sum()",
    test_type="not_regex",
    pattern=r"\bsum\s*\(",
    error_message="Implement the summation manually — do not use sum().",
)

Both types use re.MULTILINE | re.DOTALL flags.


variable — value with validator

Reads a variable from the IPython namespace and passes it to a validator callable. The test passes when validator(value) returns True.

TestCase(
    name="score is between 0 and 100",
    test_type="variable",
    variable_name="score",
    validator=lambda v: 0 <= v <= 100,
    error_message="score = {value}, but it must be between 0 and 100.",
)

Use {value} as a placeholder in error_message — it is replaced with the actual variable value when the test fails.


type_check — isinstance check

Passes when isinstance(value, expected) is True. Works with a function's return value or a named variable.

# Check a function's return type
TestCase(
    name="get_scores() returns a list",
    test_type="type_check",
    function_name="get_scores",
    inputs=[],
    expected=list,
)

# Check a variable's type
TestCase(
    name="result is a float",
    test_type="type_check",
    variable_name="result",
    expected=float,
)

# Accept multiple types using a tuple
TestCase(
    name="index is int or float",
    test_type="type_check",
    variable_name="index",
    expected=(int, float),
)

partial_output — fuzzy match

Passes when the Levenshtein similarity between the actual and expected output is ≥ similarity_threshold. Useful when minor formatting differences (extra spaces, punctuation) should still receive credit.

TestCase(
    name="Greeting (fuzzy match)",
    test_type="partial_output",
    stdin_input="Bob",
    expected="Hello, Bob! Welcome.",
    similarity_threshold=0.85,
    description="At least 85% similar to the expected greeting.",
)

similarity_threshold must be a float in (0.0, 1.0]. A value of 1.0 requires an exact match; 0.8 allows up to ~20% edit distance.


regex_output — stdout regex

Passes when re.search(pattern, output) finds a match in the printed output. Good for checking output format without requiring an exact string.

TestCase(
    name="Output contains a decimal number",
    test_type="regex_output",
    pattern=r"\d+\.\d+",
    error_message="Expected a decimal number in the output.",
)

# With a specific function
TestCase(
    name="format_price() output includes currency symbol",
    test_type="regex_output",
    function_name="format_price",
    inputs=[9.99],
    pattern=r"\$\d+\.\d{2}",
)

contains_output — stdout substring

Passes when expected is a substring of the printed output (case-sensitive).

TestCase(
    name="Output mentions 'Error'",
    test_type="contains_output",
    expected="Error",
    description="The program must print an error message.",
)

multiline_output — order-independent lines

Splits expected on newlines and verifies that every non-empty line appears somewhere in the output. Order does not matter — useful when students may print items in any sequence.

TestCase(
    name="Prints name and age",
    test_type="multiline_output",
    expected="Name: Alice\nAge: 30",
    description="Both lines must appear in the output (any order).",
)

Sections

Use TestSection to group tests into named blocks, each rendered as a separate table with its own header. Import it alongside ColabTestFramework and TestCase, then call run_sections() instead of run_tests().

from notebook_cell_tester import ColabTestFramework, TestCase, TestSection

tester = ColabTestFramework()

tester.run_sections([
    TestSection("Part 1: Code structure", [
        TestCase(
            name="Uses a for loop",
            test_type="regex",
            pattern=r"\bfor\b",
            description="Your solution must iterate with a for loop.",
        ),
        TestCase(
            name="Does not use sum()",
            test_type="not_regex",
            pattern=r"\bsum\s*\(",
            error_message="Implement the summation manually — do not call sum().",
        ),
    ]),
    TestSection("Part 2: Correctness", [
        TestCase(
            name="total([1, 2, 3]) == 6",
            test_type="return",
            function_name="total",
            inputs=[[1, 2, 3]],
            expected=6,
        ),
        TestCase(
            name="total([]) == 0",
            test_type="return",
            function_name="total",
            inputs=[[]],
            expected=0,
        ),
    ]),
])

tester.display_results()

display_results() automatically detects that sections were used and renders a named header bar above each group showing the section title and its own n/total passed count, followed by an overall summary banner at the top.

run_tests() with a flat list continues to work exactly as before — sections are opt-in.


TestCase field reference

Field Type Required for Notes
name str all Display name shown in the results table
test_type str all One of the test type strings above
function_name str function tests, type_check (one of) Looked up from get_ipython().user_ns
variable_name str variable, type_check (one of) Looked up from get_ipython().user_ns
inputs list function tests Positional arguments; defaults to []
stdin_input str output tests Simulates input() calls; lines separated by \n
expected Any most tests Return value, output string, exception type, or target type
similarity_threshold float partial_output Must be in (0.0, 1.0]
tolerance float return (optional) Passes when abs(got - expected) <= tolerance; must be >= 0
validator Callable variable lambda value: bool
pattern str regex, not_regex, regex_output Python regex string
description str optional on all Shown as an italic subtitle under the test name
error_message str optional on all Custom failure message; use {value} placeholder in variable tests

Results display

display_results() renders an HTML table inside the notebook:

  • Summary banner — gradient bar showing passed / total (%) with a congratulatory message when all tests pass.
  • Section headers (when using run_sections()) — a purple gradient bar above each group showing the section name and its own n/total passed count.
  • Status column — green ✓ PASS or red ✗ FAIL badge per test row.
  • Test column — test name, with description shown as a small italic subtitle when set.
  • Details column — human-readable message (expected vs. actual values, similarity percentages, missing lines, etc.).
  • Collapsible traceback — when an unhandled exception occurs during execution, a ▶ Show technical details disclosure element hides the raw traceback by default, keeping the table readable for beginners.

Common patterns

Test an entire cell (no function)

Omit function_name. The framework re-executes the student's full cell source in an isolated namespace with stdin redirected.

TestCase(
    name="Reads two numbers and prints their sum",
    test_type="output",
    stdin_input="3\n7",
    expected="10",
)

Combine structural and functional tests

tests = [
    # 1. Check the algorithm used
    TestCase(
        name="Uses recursion",
        test_type="regex",
        pattern=r"def\s+factorial.*?\n(?:.*\n)*?.*\bfactorial\s*\(",
        description="factorial() must call itself recursively.",
    ),
    # 2. Check correctness
    TestCase(
        name="factorial(5) == 120",
        test_type="return",
        function_name="factorial",
        inputs=[5],
        expected=120,
    ),
    # 3. Check error handling
    TestCase(
        name="Raises ValueError for negative input",
        test_type="exception",
        function_name="factorial",
        inputs=[-1],
        expected=ValueError,
    ),
]

Check a variable's value and type together

tests = [
    TestCase(
        name="result is a float",
        test_type="type_check",
        variable_name="result",
        expected=float,
    ),
    TestCase(
        name="result is positive",
        test_type="variable",
        variable_name="result",
        validator=lambda v: v > 0,
        error_message="result = {value}, but it must be greater than 0.",
    ),
]

Flexible output with partial_output

When students may phrase output slightly differently, use partial_output to award credit for near-correct answers:

TestCase(
    name="Farewell message",
    test_type="partial_output",
    stdin_input="Maria",
    expected="Goodbye, Maria! See you soon.",
    similarity_threshold=0.80,
    description="Your farewell message must be at least 80% similar to the expected output.",
)

License

MIT — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

notebook_cell_tester-1.5.0.tar.gz (23.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

notebook_cell_tester-1.5.0-py3-none-any.whl (19.9 kB view details)

Uploaded Python 3

File details

Details for the file notebook_cell_tester-1.5.0.tar.gz.

File metadata

  • Download URL: notebook_cell_tester-1.5.0.tar.gz
  • Upload date:
  • Size: 23.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for notebook_cell_tester-1.5.0.tar.gz
Algorithm Hash digest
SHA256 8edd4f2059f78c64116ddd563eab4f9b3f211b66c90eb72398537aae1d85aee9
MD5 b3b7f038fa2d7a0d8326750ab6cae812
BLAKE2b-256 378aa70314b8b99a5b41f7b0017be55fa21a014a74a23ac09b0c32691fd014f7

See more details on using hashes here.

Provenance

The following attestation bundles were made for notebook_cell_tester-1.5.0.tar.gz:

Publisher: publish.yml on EnriqueVilchezL/notebook_cell_tester

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file notebook_cell_tester-1.5.0-py3-none-any.whl.

File metadata

File hashes

Hashes for notebook_cell_tester-1.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 8058dbfae166c2edeb6bc1385af3087d0f948f291e2713d499b7e8887f591d82
MD5 24602dc396e9262ef027da74ce58fadf
BLAKE2b-256 5b1108abecf22cddd24607bd1d4d420d9134c99c26b3fbeae8d2f171b960af6b

See more details on using hashes here.

Provenance

The following attestation bundles were made for notebook_cell_tester-1.5.0-py3-none-any.whl:

Publisher: publish.yml on EnriqueVilchezL/notebook_cell_tester

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page