Русский | English
fastmem
High-performance process memory reading for Windows. Pure ctypes and the
standard library. An optional C extension is built automatically when a
compiler is available, but is never required.
Built for reverse engineering, debugging, memory forensics and security research.
Installation
pip install fastmem
From source, with a local C extension build:
git clone https://github.com/slitter-tech/fastmem.git
cd fastmem
pip install -e .
Quick start
from fastmem import Process
with Process(pid) as p:
print(p.machine_name, p.pointer_size * 8, "bit")
data = p.read(address, 32) # bytes
value = p.read(address, as_="float") # number
obj = p.read(address) # pointer_size bytes
Hot loops use or_none=True, which returns None instead of raising:
hit = p.read(address, 8, or_none=True)
API
Eight methods cover everything.
| Method | Result | Use for |
|---|---|---|
read(addr, size=None, as_=None, or_none=False) |
bytes, number or None |
single reads |
read_many(addrs, size=None, as_=None, into=False, span=-1, threads=0) |
list or bytearray |
batches |
read_region(base, size, chunks=0, or_none=False) |
bytes or generator |
scanners |
regions(min_size=0, max_size=0, ...) |
generator of Region |
virtual memory walk |
find(pattern, regions=None, align=1, limit=0, chunk_size=0) |
generator of addresses | value search |
query(addr) |
Region or None |
describe one region |
is_alive() |
bool |
liveness |
close() |
- | release the handle |
Reading
size=None means the target pointer size, so read(addr) reads a pointer
directly. as_ interprets the bytes as a number:
as_ |
Meaning |
|---|---|
None |
raw bytes |
'ptr' |
pointer, following the target bitness (8 on x64/ARM64, 4 on x86) |
'int', 'uint' |
unsigned 32-bit |
'i32', 'u32' |
signed / unsigned 32-bit |
'i64', 'u64' |
signed / unsigned 64-bit |
'i8', 'u8', 'i16', 'u16' |
narrow integers |
'float', 'f32' |
32-bit float |
'double', 'f64' |
64-bit float |
'long', 'ulong' |
unsigned 64-bit |
int and long are unsigned on purpose: memory reads are about raw values
(hashes, flags, pointers), and 0xDEADBEEF coming back as -559038737
surprises everyone. Use i32 / i64 when you want signed interpretation.
with Process(pid) as p:
hp = p.read(addr, as_="float")
n = p.read(addr, as_="int")
ok = p.read(addr, as_="ptr", or_none=True)
Batches
with Process(pid) as p:
values = p.read_many(addrs, as_="ptr") # list[int | None]
raw = p.read_many(addrs, into=True) # one bytearray, failures zeroed
span groups addresses closer than span bytes and reads each group with
a single call. -1 (default) uses four pages, 0 disables grouping.
Nearby addresses are the common case in real work - object fields, array
elements, list nodes - and grouping is up to 67x faster than reading
them one at a time.
threads spreads the addresses over worker threads and requires the C
extension. The workers live in C, are created on first use and are reused
afterwards, so the cost is paid once rather than per call. Measured gain is
1.5x at 2 threads and 2.6x at 4 on addresses too spread out for
grouping; grouping still wins whenever it applies, so pair threads with
span=0 when the addresses are scattered.
Regions
with Process(pid) as p:
for reg in p.regions(max_size=64 << 20):
blob = p.read_region(reg.base, reg.size, or_none=True)
if blob:
scan(blob)
# stream a large region instead of holding it in memory
for chunk in p.read_region(base, size, chunks=1 << 20):
scan(chunk)
Search
import struct
from fastmem import Process
needle = struct.pack("<Q", 0x1A2B3C4D5E6F7788)
with Process(pid) as p:
for addr in p.find(needle, align=8, limit=1000):
print(hex(addr))
The search runs in C through bytes.find: locating a pattern inside a
megabyte region takes ~0.3 us, while a Python loop over the bytes is five
orders of magnitude slower. chunk_size reads in overlapping blocks so
hits straddling a boundary survive.
Target properties
Detected automatically via IsWow64Process2 → IsWow64Process → system
architecture, so as_='ptr' follows the target, not the host.
p.pointer_size # 4 or 8
p.is_64bit # bool
p.machine_name # 'x86', 'x64', 'ARM64', ...
p.page_size # from GetSystemInfo
p.backend() # 'c-extension' or 'python-ctypes'
Examples
Walk a pointer chain:
with Process(pid) as p:
node = p.read(root, as_="ptr")
chain = []
while node and len(chain) < 100:
chain.append(node)
node = p.read(node, or_none=True)
if node:
node = int.from_bytes(node, "little")
print("chain length:", len(chain))
Dump every readable region:
with Process(pid) as p:
total = 0
for reg in p.regions(max_size=64 << 20):
blob = p.read_region(reg.base, reg.size, or_none=True)
if blob:
total += len(blob)
print("read", total, "bytes")
Compatibility
- Windows 7, 8, 8.1, 10, 11, Server 2012+; x86, x64, ARM64.
- Python 3.9-3.13. PyPy works on pure Python.
- 32- and 64-bit targets, including WOW64.
- Protected processes (PPL) and PID 4 (System): insufficient rights give
ProcessOpenErrorwith readable text and a Windows error code, never a crash. - A process dying mid-read:
or_none=TrueyieldsNone, plainreadraisesProcessTerminatedError(aReadMemoryErrorsubclass). - Minimum access rights by default:
PROCESS_VM_READ | PROCESS_QUERY_LIMITED_INFORMATION, which works without elevation for processes of the same user. Override withProcess(pid, access=...). - No hardcoded Windows structure offsets and no assembly: only public APIs, each checked for availability.
- Page sizes are read from
GetSystemInfo, never assumed.
Thread safety
A Process instance is not thread-safe (it holds a reusable buffer).
Open one Process per thread for manual parallel reads. read_many is the
exception: the worker pool is shared, guarded, and rebuilt on demand, so it
is safe to call from several Python threads at once.
Performance
Python 3.11, 4 cores, foreign process, 8 bytes per address. "Python" means
the same work through ctypes with the extension removed, so the columns
differ only in where the loop runs.
| Operation | Python | C | Speedup |
|---|---|---|---|
read() in try/except, bad address |
10.46 | - | - |
read(or_none=True), bad address |
3.55 | - | - |
read() success |
3.53 | - | - |
read(as_='int') |
5.25 | - | - |
batch into bytearray |
5.08 | 1.40 | 3.62x |
batch into list[bytes] |
5.82 | 1.36 | 4.27x |
batch into list[int] |
5.29 | 1.47 | 3.59x |
| grouping vs streaming | 5.12 | 0.076 | 67.2x |
Large blocks, one API call each:
| us | MB/s | |
|---|---|---|
read_region(4 KiB) |
9.7 | 421 |
read_region(64 KiB) |
22.0 | 2977 |
read_region(1 MiB) |
817 | 1284 |
Where the speed comes from
- Block size matters more than language. One
read_region(64 KiB)costs ~22 us; reading the same 64 KiB address by address costs ~150 ms. - Search belongs in C.
bytes.findbeats a Python loop by five orders of magnitude. - Allocation zeroes memory.
(ctypes.c_char * n)()memsets: ~400 us for 1 MiB. A reusable buffer removes it. - Exceptions are expensive. Formatting the error text cost 3.3 us; the
message is now built lazily in
__str__. - C means one boundary crossing instead of a thousand. A Python -> C call costs ~0.6 us, and a 1000-address batch in Python pays that a thousand times. The C loop crosses once.
- Handing data to C has a price too.
(ctypes.c_size_t * n)(*addrs)took 330 us for 2000 addresses;array.arrayfills the same buffer in 66 us. Worth 12% of a threaded read on its own.
Threads
Measured on a foreign process, no grouping, microseconds per address:
| 1 thread | 2 | 4 | |
|---|---|---|---|
| C, worker pool | 1.28 | 0.85 (1.51x) | 0.50 (2.57x) |
| same, into a bytearray | 1.26 | 0.96 (1.32x) | 0.54 (2.32x) |
| same, sparse addresses | 1.24 | 0.93 (1.33x) | 0.69 (1.80x) |
Threads used to be a net loss here: building threading.Thread objects per
call cost ~812 us on Windows, more than the ~1700 us of work it was meant
to parallelise, so four threads measured 1.11x on dense input and 0.20x on
sparse input - worse than not threading at all. The workers now live in C,
are created on first use and park on an event between calls, which is what
turned those into 2.57x and 1.80x. Repeated runs put the four-thread figure
between 2.57x and 2.71x; absolute timings move with machine load, the
ratios do not.
A control group confirms the mechanism: a variant that holds the GIL gains exactly nothing from four threads (1617 us against 1605 us for one), so the limiter is the GIL, not contention inside the kernel. The pool releases it for the whole read.
The calling thread takes a share of the work, so threads=4 means five
readers. threads is capped at len(addrs) - 1 for the same reason.
Prefer grouping for dense addresses: one read per window beats one read per address by more than threads can recover. Reach for threads when the addresses are too spread out for grouping to apply.
Building from source
python setup.py build_ext --inplace # optional
python -m tests.test_fastmem # functional tests
python -m tests.benchmarks # benchmarks, own process
python -m tests.benchmarks --foreign # benchmarks, foreign process
Building the extension needs MSVC and the Windows SDK. The SDK is not
always in the default location, so setup.py looks through WindowsSdkDir
and the usual paths including G:\Windows Kits\10. Some Python builds ship
no pythonXY.lib (built without Py_ENABLE_SHARED); in that case the
import library is generated from the DLL export table with dumpbin and
lib.exe.
License
MIT. See LICENSE.
Metadata
Release files for fastmem 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| fastmem-0.1.0.tar.gz | 60.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| fastmem-0.1.0-cp311-cp311-win_amd64.whl | CPython 3.11 | CPython 3.11 | Windows x86-64 | Details |
Total release size: 110.0 kB
Release files / fastmem-0.1.0.tar.gz
| Download URL | fastmem-0.1.0.tar.gz |
|---|---|
| Size | 60.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
03b79467e75fd7989b7d4fe1bb3242b4c23cd59ee272e8c9cd7986ab32cd60d0
|
|
BLAKE2b-256 checksum How to use checksums |
9956af65d5931cd05b4743dd5171270e4c8ee81d7641f19f9e7efbc3c70d87a2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency logRelease files / fastmem-0.1.0-cp311-cp311-win_amd64.whl
| Download URL | fastmem-0.1.0-cp311-cp311-win_amd64.whl |
|---|---|
| Size | 49.6 kB |
| Tags | CPython 3.11 Windows x86-64 |
|
SHA-256 checksum How to use checksums |
20d5d8fb2f29c953508923ff7f7656321020703059ab620e45ee7dcc7229cecc
|
|
BLAKE2b-256 checksum How to use checksums |
1cc1d05cbf38918b65b6107dfc174c7d17e052a5b7a8a62f95f8859c4f1fb665
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency log