Skip to main content

setalign

setalign provides a toolkit for analyzing citation alignment between agents (e.g., courts, institutions, documents) using information-theoretic principles.

🎯 Motivation

The idea for this project arose from the analytical need of comparing citations between institutional actors. As readers of judicial politics will know, politics happens through law, and citations to caselaw has been widely used to scaled judical opinions (see works by Jeffrey Lax, Jonatha Kastellec, Erik Voeten, Daniel Naurin for instance). However, a uncharted territory is how to robustly compare citation set's between agents when you expect them to be very similar, if not identical. At face value, this seems like a non-issue all that you do is count overlaps and that will easily tell you the degree of precisely, overlap. Yet, what about when the second mover overlaps entirely with the first one but also adds an extra citation? That should not be scaled equally as one with a perfect one-to-one overlap, at least not so if we are saying that legal ciations matter.

AGENT_1 = {A,B,C,D} AGENT_2 = {Z,Y,X,W}

AGENT_1 = {A,B,C,D} AGENT_2 = {A,B,C,D,E}

🧮 Core Methodology

This module integrates two methodological innovations:

  1. Set Citation Alignment Ratio (SCAR) — measures informational agreement between citation sets.
  2. Locally Information-Adjusted Citations (LIAC) — transforms SCAR into interpretable, scalable alignment units.

Method Overview

  1. Applies information penalties for both missing and extra citations.
  2. Scales by local citation universe size to enable cross-case comparability.
  3. Produces bounded, interpretable measures ranging from –γ to +γ.

🧠 Primary Use Cases

  • Judicial citation behavior (ideological alignment, institutional influence)
  • Legislative or policy consensus measurement
  • Academic citation analysis (political science, law, etc.)
  • Cross-institutional alignment studies
  • Policy or regulatory document concordance

⚙️ Walkthrough With Example Data

The example dataset contains paired citation lists for two actors. Each row is a dyad-level comparison.

from setalign import (
    calculate_liac,
    calculate_scar,
    calculate_scar_metrics,
    extract_citation_sets,
    load_ecj_example_data,
    scale_scar_by,
)

# Load the packaged ECJ example data
ecj_data = load_ecj_example_data()

The citation columns must contain list-like citation sets. In this example:

  • actor_a_citations is the reference or agenda-setting actor.
  • actor_b_citations is the responding actor.
  • dyad_id identifies the paired comparison.

The packaged example dataset is intentionally minimal. It contains only the columns used in this walkthrough: dyad_id, actor_a_citations, and actor_b_citations.

Calculate SCAR

calculate_scar() compares the two citation lists and adds:

  • case_flag: the set relationship case.
  • scar: the Set Citation Alignment Ratio.

Semantically, SCAR asks how much informational agreement exists between the reference actor's citation set and the responding actor's citation set. It is not only a raw overlap count. Missing citations and additional citations both matter because each changes the information contained in the second actor's response.

para_level_data = calculate_scar(
    ecj_data,
    input_col1="actor_a_citations",
    input_col2="actor_b_citations",
)

Add Supporting Citation Metrics

calculate_scar_metrics() adds interpretable set counts, including the intersection size, overall difference, initial citation universe size, and joint citation universe size.

These columns make the SCAR score auditable. For example, intersection tells you how many citations both actors share, while overall_difference tells you how much citation material appears in only one actor's set. init_citation_universe_len is especially important because it records how much citation information actor A placed on the table.

para_level_data = calculate_scar_metrics(
    para_level_data,
    input_col1="actor_a_citations",
    input_col2="actor_b_citations",
)

Extract Citation Sets

extract_citation_sets() adds the actual citation sets behind the metrics:

  • shared_citations_list
  • agent1_unique_citations_list
  • agent2_unique_citations_list
  • universe_citations_list

This step is useful when you want to move from measurement back to interpretation. Instead of only knowing that two actors diverged, you can inspect which citations were shared, which were unique to actor A, and which were introduced by actor B.

para_level_data = extract_citation_sets(
    para_level_data,
    input_col1="actor_a_citations",
    input_col2="actor_b_citations",
)

Calculate LIAC

calculate_liac() transforms SCAR into Locally Information-Adjusted Citations. The default scaling term is the initial citation universe size, here stored in init_citation_universe_len.

The intuition is local comparability. A SCAR score of 0.50 does not mean the same thing when actor A cited one authority as when actor A cited twenty. LIAC scales the alignment score by the size of actor A's citation universe, so the final value reflects both alignment quality and the amount of citation information at stake in that dyad.

para_level_data["liac"] = calculate_liac(
    para_level_data,
    init_citation_universe_len="init_citation_universe_len",
)

You can also scale SCAR by a non-local value using scale_scar_by(). This is useful when you have theoretical priors about the citation universe that should anchor comparisons across rows. For example, a global median scaling term treats the typical citation-universe size as the reference point, rather than letting each dyad define its own scale. This is less naive than choosing an arbitrary constant, but it is still a modeling choice: use it when the global reference value has a substantive interpretation for your application.

para_level_data["liac_global_median"] = scale_scar_by(
    para_level_data,
    scar_col="scar",
    scaling_term=para_level_data["init_citation_universe_len"].median(),
)

Inspect Results

result_cols = [
    "dyad_id",
    "case_flag",
    "scar",
    "intersection",
    "overall_difference",
    "init_citation_universe_len",
    "joint_citation_universe_len",
    "liac",
    "liac_global_median",
]

print(para_level_data[result_cols].head(8))

Example output, shown in a compact -style layout:

Rows: 8
Columns: 9

dyad_id                       DYAD001, DYAD002, DYAD003, DYAD004, DYAD005, DYAD006, DYAD007, DYAD008
case_flag                     2, 1, 1, 1, 1, 2, 2, 3
scar                          0.250, 0.333, 0.333, 0.333, 1.000, 0.333, 0.333, 0.019
intersection                  1, 1, 1, 1, 1, 1, 1, 0
overall_difference            3, 2, 2, 2, 0, 2, 2, 2
init_citation_universe_len    2, 1, 1, 1, 1, 2, 2, 1
joint_citation_universe_len   4, 3, 3, 3, 1, 3, 3, 2
liac                          -1.000, -0.334, -0.334, -0.334, 1.000, -0.668, -0.668, -0.962
liac_global_median            -0.500, -0.334, -0.334, -0.334, 1.000, -0.334, -0.334, -0.962

🎓 Citation

If you use this framework in academic research, please cite:

Mandujano Manríquez, M. (2026). setalign: Information-Theoretic Citation Alignment Module. Github: https://github.com/mauriciomm7/setalign

@misc{mandujano2026setalign,
  author       = {Mauricio Mandujano Manríquez},
  title        = {setalign: Information-Theoretic Citation Alignment Module},
  year         = {2026},
  howpublished = {\url{https://github.com/mauriciomm7/setalign}},
  note         = {GitHub repository}
}

📄 License

This project is licensed under the MIT License.

✅ TODO

  • UPDATE caculate_scar such that it lets you decide who is the "referenece" agent A or B or create equivalent function but flip case one to zeroes.

READ References

Information Theory:

  • Shannon, Claude E. (1948). "A Mathematical Theory of Communication." Bell System Technical Journal, 27(3), 379-423.
  • Leydesdorff, Loet. Various works on information-theoretic approaches to citation analysis - His theoretical framework treats citations as selections operating in complex information systems

Bibliometrics:

  • Lam, Weng Hoe, et al. (2022). "Bibliometric Analysis of Information Theoretic Studies." Entropy, 24(10), 1359 - Shows the growing application of information theory across disciplines
  • Garfield, Eugene (1996). Multiple works on citation analysis fundamentals and strategic aspects
  • Wouters, Paul (1997). "Citation cultures" - Addresses how citation practices vary and can be strategic
  • Validation of Google Scholar as an Impact Measure for Political Science
  • Hill, K.Q. (2022). "Web of Science Book Citation Indices and the Representation of Political Science." Journal of Electronic Publishing - Addresses unique citation challenges in political science.

Mathematical Framework:

  • Jaccard, Paul (1912) - For set similarity measures (Jaccard coefficient) as a baseline comparison

Metadata

Release files for setalign 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for setalign 0.1.0
File Size Uploaded
setalign-0.1.0.tar.gz 298.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for setalign 0.1.0
File Interpreter ABI Platform
setalign-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 590.9 kB

Release files / setalign-0.1.0.tar.gz

Download URL setalign-0.1.0.tar.gz
Size 298.0 kB
Tags Source
SHA-256 checksum
How to use checksums
c949569d6ee7395ff52d289ed4defe00e78d085c07343d582035f8905747cfa1
BLAKE2b-256 checksum
How to use checksums
a24ee7f545d8248c4fda1dc263a6170ef26695e0d450640686d4f9fcdc5efe04
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.7

Release files / setalign-0.1.0-py3-none-any.whl

Download URL setalign-0.1.0-py3-none-any.whl
Size 292.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7c234208360999978018f6235f0f67aaa06f10d90b08cd712305657345562c67
BLAKE2b-256 checksum
How to use checksums
0489bb9125e371199df0f6abfa300d23a35e5e15b88bcf50f8ce4b2c96fd6a25
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.7

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page