Skip to main content

sws

PyPI - Version Tests codecov PyPI - License

Minimal, predictable, footgun-free configuration for deep learning experiments. The most similar existing ones are OmegaConf and ConfigDict - if you are happy with them, you probably don't need this. If you want some lore, have a look at the end.

The remainder of this readme follows the CODE THEN EXPLAIN layout. The example/ folder contains a nearly real-world example of structuring a project. Install instructions at the end.

Basics

from sws import Config

# Create the config and populate the fields with defaults
c = Config()
c.lr = 3e-4

# Alternative shorthand handy for very small configs:
c = Config(lr=3e-4)

# How to make a field depend on others?
c.wd = c.lr * 0.1  # ERROR: c is write-only.
# Instead, use a lambda to make the value "lazy"
c.wd = lambda: c.lr * 0.1

# Finalizing resolves all fields to plain values, and integrates CLI args:
c = c.finalize(argv=sys.argv[1:])
assert c.lr == 3e-4 and c.wd == 3e-5

train_agi(lr=c.lr, wd=c.wd)

sws clearly separates two phases: config creation, and config use. At creation time, you build a (possibly nested) Config object. To avoid subtle bugs common in many config libraries I've used before, at creation time, the Config object is write-only; you cannot read its values. Once you finished building it up, a call to c.finalize() turns it into a read-only FinalConfig object that contains "final" values for all fields.

This finalization step can also integrate overrides from, for example, commandline arguments; more on that a little later. You can call finalize(argv) repeatedly; each call starts from the builder's original values. Finalization state is isolated per call, so calls on the same builder may also run concurrently. Container graphs in each result are defensively copied: they remain ordinary Python values, but mutating those containers cannot affect the builder or another finalization result. Opaque non-container values retain their identity.

If you want to make one field's value depend on another field's value, you can do so by wrapping the value in a lambda, which computes the derived value. This lambda will be called during finalization, where concrete config values can be accessed. In this way, in the example above, the wd setting will use the correct value of c.lr even when it is overridden by commandline arguments during finalize. This works transitively, just as you'd expect it to. Each lazy field is evaluated at most once per call to finalize; repeated dependency reads reuse that field's resolved value.

Since callable values receive this special treatment, if you want to actually set a config field's value to an actual function, that needs to be wrapped by sws.Fn:

from sws import Fn

# If you want to store a callable as a value (not execute it at finalize), wrap it:
c.log_fn = Fn(lambda s: print(s))
c = c.finalize()

# Five moments later...
c.log_fn("After finalization, the config field is just this plain function")

Nesting

Of course any respectable config library allows nested structures:

from sws import Config

# Create the config and populate the fields with defaults
c = Config()
c.lr = 3e-4
c.model.depth = 4  # No need to create parents first.

# In a nested field, lazy and `c` work just as you'd expect them to:
c.model.width = lambda: c.model.depth * 64
c.model.emb_lr = lambda: c.lr * 10 / c.model.width

c = c.finalize()

# Pass model settings as kwargs, for example:
m = MyAGIModel(**c.model.to_dict())
train_agi(m, c.lr)

The reason we need to_dict() above is that FinalConfig implements as few methods as possible, to leave as many names as possible free to be used for configs. For instance, keys, values, and items are not implemented so that you can use them as config names. This also means, that it doesn't implement the Mapping protocol and can't be **'ed. So, just call to_dict, it's fine.

You don't really need to know this, but internally, the full config is stored as a flat dict ("model.emb_lr" is a key), and subfields are just prefix-views into that dict.

Commandline overrides

The finalize() method allows you to pass a list of argv strings to it that serve as overrides:

from sws import Config

c = Config(lr=1.0, model={"width": 128, "depth": 4})
c = c.finalize(["c.model.width=512", "c.model.depth=2+2"])

# However, we're lazy. The shortest unique segment suffix works:
c = c.finalize(["width=512", "depth=2+2"])

# In real life, you'd probably pass sys.argv[1:] instead.

Only the syntax a=b is supported (not a b or --a b), any argument without = is ignored by finalize (it is returned as unused when you pass return_unused_argv=True). Note that sws.run, described below, is stricter: it raises on such leftover arguments unless you use forward_extras=True. This is to reduce ambiguity and allow catching typos.

The values of the overrides are parsed as Python expressions using the simpleeval library. This makes a lot of Python code just work, for example you can write model.vocab=[i*i for i in range(10)] and it'll work. You can also access the current config using the name c, so something like 'c.model.width=3 * c.model.depth' works. Note that I quoted the whole thing, for two reasons: (1) to stop my shell from interpreting * as wildcard, and (2) because I used spaces.

At the same time, string values just work without quoting: a value that is not valid Python (msg=hello world, path=/data/foo) or consists only of unknown bare words (dataset=imagenet_2012, arch=gpt-4) is taken as a literal string. Anything else that fails to evaluate is an error, not a string: broken expressions (lr=1/0), unknown functions, and anything mentioning c (so wd=c.lrr * 0.1 reports the typo instead of silently assigning a string). The common typos true/false/none/null error with a hint towards the Python spelling. To force a string that would otherwise evaluate, quote it for Python too: name="'True'". Expressions referencing c see the final config values, after all overrides are applied — including overrides that appear later in the argument list — so the result does not depend on the order of the arguments. After an override key is resolved and its value is evaluated, the value is assigned with the same shape rules as config construction: dicts create subtrees, leaves replace groups, and groups replace leaves. One caveat: override values are only evaluated at the end of finalization, so the children of a dict-valued override (like model=dict(width=64)) do not exist yet while the remaining arguments are processed. Targeting them in the same argv (like a subsequent model.width=128) is therefore an error; adjust the dict expression itself instead. (A subsequent model.width:=128 follows the usual shape rules: creating that exact key replaces the model leaf wholesale.)

For convenience, the keyname can be shortened to the shortest unique suffix across the whole config (i.e. all nesting levels). For example, model.head.lr can be shortened to head.lr or lr if unambiguous. In the case of ambiguity, sws errs on the cautious side and errors out. You can always specify the full name starting with c. to be perfectly unambiguous. Invalid, unknown, and ambiguous override keys raise sws.OverrideError, a subclass of sws.FinalizeError.

If there's a name that you use many times, and you'd like to set all matching keys to a specific value, use the wildcard prefix syntax ..name=value. For example, if c.head.lr and c.body.lr both exist, you may use ..lr=0.1 to set both simultaneously. Note that this is a "plaintext" wildcard, so it will also match c.flip_lr. If you want to match only full leaf names, just add a dot: ...lr=0.1, since this matches the suffix .lr.

Finally, the syntax name:=value creates the exact field c.name even if it does not exist. This can be useful when the codebase uses the pattern c.get("name", default) for things, and the get_config doesn't include a value for name. Use with care though. The := marker is only recognized between the key and value, so normal override values may contain := as plain text.

sws.run and suggested code structure

The train.py file could look something like this:

import sws

# ...lots of code...

def train(c):
    # Do some AGI things, but be careful please.
    # `c` is a FinalConfig here, i.e. it's been finalized.

if __name__ == "__main__":
    sws.run(train)

This seemingly innocuous code does a lot, thanks to judiciously chosen default arguments. The full call would be sws.run(train, argv=sys.argv[1:], config_flag="--config", default_func="get_config").

First, it looks for a commandline argument --config filename.py (or --config=filename.py).

It then loads said file, and runs the get_config function defined therein, which should return a fully populated sws.Config object. Note that it's plain python code, so it may import things, have a lot of logic, feel free to do as much or as little as you want.

Finally, it finalizes the config with the remaining commandline arguments, and calls the specified function (in this example, train) with the FinalConfig.

Here's what a config file might look like, let's call it vit_i1k.py:

from sws import Config

def get_config():
    c = Config()
    c.lr = 3e-4
    c.wd = lambda: c.lr * 0.1
    c.model.name = "vit"
    c.model.depth = 8
    c.model.width = 512
    c.model.patch_size = (16, 16)
    c.dataset = "imagenet_2012"
    c.batch = 4096
    return c

Then, you would run training as python -m train --config vit_i1k.py batch=1024. In a real codebase, you'd have quite a few config files, maybe in some structured config/ folder with sub-folders per project, user, topic, ...

There's three more things sws.run does for convenience:

  • If no --config is passed, it looks for the get_config function in the file which called it. This is very convenient for quick small scripts. Two caveats: the file is re-executed to find that function, so all of its top-level code runs a second time (keep side effects under the if __name__ == "__main__": guard); and "the file which called it" is the direct caller, so if you wrap sws.run in a helper function of your own, it will look in your helper's file instead — pass --config explicitly then.
  • If you use run(fn, forward_extras=True), then all unused commandline arguments, i.e. all those without a =, are passed in a list as the second argument to fn. This can be used to do further custom processing unrelated to sws. If forward_extras is False and any such extra tokens are present, sws.run raises a ValueError listing the unused arguments.
  • For extra flexibility, you can actually specify which function should be called. The syntax is --config file.py:function_name, it's just that the function name defaults to get_config. This way, you can have multiple slight variants in the same file, for example.

See the example/ folder of this repo for a semi-realistic example, including a sweep to run sweeps.

A realistic example

This is how I'd structure a codebase, roughly. See also example/ folder.

Various experiment configurations in the configs/ folder. For example, configs/super_agi.py:

from sws import Config

def get_config():
    c = Config()
    c.lr = 0.001
    c.wd = lambda: c.lr * 0.1
    c.model.depth = 4
    c.model.width = 256
    c.model.heads = lambda: 4 if c.model.width > 128 else 1
    return c

Your main code, for example train.py:

from sws import run

def main(c):
    print("Training with config:\n" + str(c))
    # Your training code here...

if __name__ == "__main__":
    run(main)

Run a different config file and override values from CLI if wanted:

python -m train --config configs/super_agi.py model.depth=32

See example/sweep.fish for a trivial sweep over a few values.

Reusable subtrees

As projects and configs grow, you may want to write helper functions to populate subtrees. The sws-blessed way to do so, which ensures that all features work as expected without footguns, is creating the subtree "in-place" as follows:

def make_tokenizer(c, ctok):
    ctok.path = lambda: f"/foo/bar/{c.voc}" if c.voc != "magic" else "/the/magic"
    ctok.regex = r"\d+" if c.voc == "magic" else "default"

c = Config()
c.voc = "not magic 123"
make_tokenizer(c, c.data.tokenizer)

This keeps all leaves visible before finalization, so everything you'd expect works: normal overrides like voc=magic and data.tokenizer.regex=r"\w+", and even creating explicit extra leaves such as data.tokenizer.special:=42.

Depending on your background, you may have defaulted to the following construction, which does not work and is not possible for sws to support without dangerous footguns:

def make_tokenizer():  # DON'T
    c = Config()
    c.path = "/foo/bar/tokenizer.model"
    c.regex = lambda: rf"{c.path}:\d+"
    return c

c = Config()
c.data.tokenizer = make_tokenizer()  # NOT RIGHT
c.data.tokenizer = lambda: make_tokenizer()  # NOT RIGHT EITHER

Since this is the first intuition for some people, sws detects this pattern and gives an error message hinting to the blessed way.

Copying an existing subtree view, such as c.model2 = c.model1, is supported only when all fields in the source subtree are eager values. sws rejects the assignment if the source contains lazy fields: their Python closures would still refer to the original subtree, so the copy could silently compute wrong values. Populate both destination views in place instead. Assigning an empty subtree view is also rejected because it has no fields to copy.

Note that a function returning plain python dictionaries works, since dictionaries are valid config leaf values, but that will not create a subtree from the dict.

Some more misc notes

  • The FinalConfig has a nice pretty printer when cast to string or printed. Fields that were set from argv are annotated with (argv) — or with (argv, as string) when the value was taken as a literal string — so a glance at the printed config shows what the CLI changed and how it was read.
  • When a dict is assigned to a Config field, it's turned into a Config.
  • Assigning a value to a group replaces its subtree (e.g. c.model = "vit" clears all c.model.*), and assigning a dict to a leaf replaces the leaf with a group.
  • Cycles in computed callables are detected and raise an exception at finalize.
  • Other exceptions raised by a lazy value are wrapped in sws.FinalizeError with the failing field's name; the original exception is available as __cause__.
  • The FinalConfig has a .to_json() and .to_flat_json() utils that return a string that's the json serialized config, but with non-json-serializable values replaced by an explanatory string. It's for logging/storing of configs for humans.
  • Similarly, there's the sws.from_json and sws.from_flat_json counterparts, they are provided purely for human analysis and convenience, since json is lossy wrt sws.

Installing

pip install sws-config

Testing

python -m pytest

TODOs

  • When passing commandline args, using lazy/lambda makes no more sense. So we should lift the requirement for Fn-wrapping of callables here. 'log_fn=Fn(lambda s: print(f"Log: {s}"))'.

Probably overkill:

  • Auto-generate a commandline --help?
  • Auto-generate a terminal UI to browse/change config values on finalize() could be fun.

Lore

You obviously wonder "Why yet another config library, ffs?!" - and you're right. There are many, but there's none that fully pleases me. So I gave in.

I've heavily used, and hence been influenced by, many config systems in the past. Most notably ml_collections.ConfigDict and chz, both of which I generally liked, but both had quite some pitfalls after serious use, which I try to avoid here. Notable examples which I used but did not like are gin, yaml / Hydra, kauldron.konfig; they are too heavy, unpythonic, and magic; there be footguns. fiddle requires your config to import everything, which I don't like. I refuse to build around types in Python, like pydantic, tyro, dataclasses, ..., so not even linking them. Finally, I haven't used, but thoroughly read Pydra and Cue, which together inspired the two-step approach with finalization.

Why is it called sws? It's a nod to OpenAI's chz config library, and the author being a very fond resident of Switzerland.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sws_config-0.10.0.tar.gz (38.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

sws_config-0.10.0-py3-none-any.whl (28.7 kB view details)

Uploaded Python 3

File details

Details for the file sws_config-0.10.0.tar.gz.

File metadata

  • Download URL: sws_config-0.10.0.tar.gz
  • Upload date:
  • Size: 38.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for sws_config-0.10.0.tar.gz
Algorithm Hash digest
SHA256 a36522c292a7c60c67818d67c9efb869fc18d5a69702d7b57cafc3a260112170
MD5 9a4261227119ccd3c9745a916c38a366
BLAKE2b-256 ebdad4803db541dce6c1634d727e632cfacbe2094d6d35015bd23c5d658170bc

See more details on using hashes here.

Provenance

The following attestation bundles were made for sws_config-0.10.0.tar.gz:

Publisher: publish.yml on lucasb-eyer/sws

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file sws_config-0.10.0-py3-none-any.whl.

File metadata

  • Download URL: sws_config-0.10.0-py3-none-any.whl
  • Upload date:
  • Size: 28.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for sws_config-0.10.0-py3-none-any.whl
Algorithm Hash digest
SHA256 d1b0a235d700c778791657d16e139fb8dc05884aab1aa16aa02fa99e46d96c35
MD5 66ab087dab9834a7245cf7b843fd467d
BLAKE2b-256 2aac4f1573b5d38da8712b0c3c1ce14f5d1b76caa8d425ea7f2ec13c2ff61332

See more details on using hashes here.

Provenance

The following attestation bundles were made for sws_config-0.10.0-py3-none-any.whl:

Publisher: publish.yml on lucasb-eyer/sws

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.10.0 This release

2 files

0.9.2

2 files

0.9.1

2 files

0.9.0

2 files

0.8.2

2 files

0.8.1

2 files

0.8.0

2 files

0.7.1

2 files

0.7.0

2 files

0.6.4

2 files

0.6.3

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.1

2 files

0.4.0

2 files

0.3.2

2 files

0.3.1

2 files

0.3.0

2 files

0.2.1

2 files

0.2.0

2 files

0.1.5

2 files

0.1.4

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page