Skip to main content
ⓘ

Downloads Downloads Coverage Status Lines of code Hits-of-Code Test-Package Python versions PyPI version Checked with mypy Ruff DeepWiki

Autofission

Autofission dynamically calculates and updates the maximum replica limit (MaxScale) for explicitly opted-in Fission Functions. It estimates the limit from currently available CPU, memory, and Pod slots on schedulable Kubernetes nodes.

A Fission Function using the newdeploy executor can scale down when demand disappears. Its Horizontal Pod Autoscaler (HPA) still needs a fixed positive MaxScale. A limit sized for today's cluster becomes too low when nodes are added. An arbitrarily high limit can flood the scheduler with Pods that cannot fit.

Autofission keeps the limit current by estimating available capacity on each node. It subtracts the requests of existing workloads and accounts for the Function's current replicas. It is designed for elastic, bare-metal, homelab, and edge clusters where nodes come and go and idle compute should remain available to Functions without scheduler preemption of existing services.

The only scaling setting Autofission changes is this upper bound; it also records the calculation in annotations. Fission's HPA and idle reaper still decide when the replica count grows and shrinks.

Table of Contents

Installation

Use the Python CLI for optional local or one-shot runs. Install the Helm chart to run Autofission continuously in a cluster.

Python CLI

For local use, Autofission requires Python 3.8 or newer. CI tests Python 3.8 through 3.14, free-threaded Python 3.14, and Python 3.15 beta:

python -m pip install autofission
autofission --help

With no credential-selection flags, the CLI uses the current kubeconfig context unless KUBERNETES_SERVICE_HOST is present, in which case it uses mounted ServiceAccount credentials. --in-cluster selects the ServiceAccount explicitly.

To run one reconciliation without starting a daemon:

autofission --once --context my-cluster

This performs a real reconciliation and can update opted-in Functions.

Install in a cluster

Prerequisites are Kubernetes and an existing Fission installation. Autofission manages only Functions that use the newdeploy executor. By default, the chart installs the controller, its RBAC, and two PriorityClasses; it does not install or remove Fission.

Set AUTOFISSION_VERSION to the chart version you want to install, then install Autofission from its OCI release:

helm upgrade --install autofission \
  oci://ghcr.io/pomponchik/charts/autofission \
  --version "${AUTOFISSION_VERSION}" \
  --namespace fission \
  --create-namespace \
  --atomic \
  --wait

For an unreleased checkout, build a controller image that the cluster can pull, replace the OCI URL with ./deploy/helm/autofission, and set its repository and tag. Autofission assumes fetcher requests of 10m CPU and 16Mi memory. If Fission uses different values, pass them to the chart so the capacity calculation remains accurate:

helm upgrade --install autofission ./deploy/helm/autofission \
  --namespace fission \
  --set-string image.repository=REGISTRY/autofission \
  --set-string image.tag=TAG \
  --set controller.fetcherCpuRequest=20m \
  --set controller.fetcherMemoryRequest=32Mi

After the chart creates its PriorityClasses, configure Fission to use the low, non-preempting runtime class. Its default value of -10 places unprioritized Pods, whose priority is normally 0, ahead of elastic Function Pods in the scheduling queue. With the default Autofission release name, add the following to Fission's Helm values and upgrade Fission before opting in any Functions:

runtimePodSpec:
  enabled: true
  podSpec:
    priorityClassName: autofission-runtime

Quick start

The example uses a Function named hello in the default namespace and assumes Fission's default same-namespace workload placement. Replace the Function name and namespace, then opt it in. If Fission sets a separate functionNamespace, use that workload namespace when querying the HPA:

kubectl label function hello \
  --namespace default \
  autoscaling.fission.io/cluster-capacity=true

After the next successful reconciliation, inspect the Function's MaxScale, the recorded calculation, and the HPA. The daemon waits 15 seconds between cycles by default; processing and transient failures can add delay:

kubectl get function hello --namespace default \
  -o jsonpath='{.spec.InvokeStrategy.ExecutionStrategy.MaxScale}{"\n"}'
kubectl get function hello --namespace default \
  -o jsonpath='{.metadata.annotations.autoscaling\.fission\.io/calculated-maxscale}{"\n"}'
kubectl get hpa --namespace default

After a successful cycle, the first two values should match. The HPA may reflect the new limit slightly later because Fission updates it asynchronously.

Removing the label stops future management. Autofission deliberately does not guess or restore a previous manually configured MaxScale; the last calculated value and informational annotations remain until you change them.

How it works

flowchart TD
    state["Cluster state<br/>Nodes, Pods, and resource requests"] --> autofission["Autofission<br/>estimates modeled capacity"]
    autofission -->|updates| limit["Function MaxScale"]
    demand["Demand or idle time"] --> scaling["Fission HPA<br/>and idle reaper"]
    limit -->|sets upper bound| scaling
    scaling -->|changes| replicas["Function replicas"]

Each cycle validates its inputs and is idempotent for an unchanged cluster snapshot:

  1. List opted-in Functions, along with Fission Environments, Kubernetes Nodes, and Pods.
  2. Keep Ready, uncordoned, non-deleting nodes. Nodes with NoSchedule or NoExecute taints are excluded by default.
  3. Calculate effective CPU and memory requests for every active, scheduled Pod, including containers, init containers and sidecars, Pod-level requests, and overhead. Exclude completed and unbound Pods.
  4. Resolve each Function's CPU and memory requests, inheriting missing or zero values from its Environment, then add the fetcher request. Use larger values observed on an existing Function Pod.
  5. Add the Function's existing Pods back to its capacity budget because step 3 counted them as other workload. This changes only the calculation, not the Pods. Then calculate how many identical Pods fit on each node. Per-node results are summed, so CPU on one node cannot combine with memory on another.
  6. Set MaxScale to at least MinScale and 1. Patch a Function when either MaxScale or Autofission's calculation annotations have changed. A resourceVersion conflict prevents concurrent edits from being overwritten.

An invalid or conflicting Function does not block the others. The controller becomes Ready after a cycle in which every managed Function completes cleanly. A later failure does not clear readiness immediately: the last successful marker remains valid until the configured probe age expires (60 seconds by default).

Configuration

CLI flags take precedence over non-empty environment variables, which take precedence over defaults.

CLI flag Environment variable Helm value Default
--interval-seconds AUTOFISSION_INTERVAL_SECONDS controller.intervalSeconds 15
--request-timeout-seconds AUTOFISSION_REQUEST_TIMEOUT_SECONDS controller.requestTimeoutSeconds 10
--retry-attempts AUTOFISSION_RETRY_ATTEMPTS controller.retryAttempts 3
--fetcher-cpu-request AUTOFISSION_FETCHER_CPU_REQUEST controller.fetcherCpuRequest 10m
--fetcher-memory-request AUTOFISSION_FETCHER_MEMORY_REQUEST controller.fetcherMemoryRequest 16Mi
--managed-label AUTOFISSION_MANAGED_LABEL controller.managedLabel autoscaling.fission.io/cluster-capacity
--managed-value AUTOFISSION_MANAGED_VALUE controller.managedValue true
--include-tainted-nodes AUTOFISSION_INCLUDE_TAINTED_NODES controller.includeTaintedNodes false
--log-level AUTOFISSION_LOG_LEVEL controller.logLevel INFO
--state-directory AUTOFISSION_STATE_DIRECTORY — /tmp/autofission (CLI); /var/run/autofission (chart)

--kubeconfig, --context, and --in-cluster select credentials. --once runs exactly one pass. --probe readiness|liveness --max-age-seconds N is intended for Kubernetes exec probes.

RBAC and security

With the default RBAC values, the chart-created ClusterRole grants only these cluster-wide operations:

  • list Nodes and Pods;
  • list Fission Environments;
  • list and patch Fission Functions.

The chart does not grant permission to read Secrets, create or delete Functions, or mutate Pods, Nodes, Deployments, or Services. Other bindings attached to the same ServiceAccount can grant additional permissions. Kubernetes RBAC cannot restrict list or patch permissions by label, so the opt-in label is an application-level boundary. Listing Pods exposes their specifications, including literal environment-variable values, to the controller process. This access is needed to protect capacity requested by other workloads.

By default, the container runs as UID/GID 65532 with a read-only root filesystem. It drops all Linux capabilities, blocks privilege escalation, and uses a RuntimeDefault seccomp profile. The chart creates a NetworkPolicy with an empty ingress list and no egress policy. Enforcement requires a compatible network plugin, and other NetworkPolicies can add allowed ingress because Kubernetes combines their rules.

The controller has a high, non-preempting PriorityClass, so a pending controller Pod is queued ahead of lower-priority Pods when capacity becomes available. It cannot evict running Pods, so the class does not guarantee availability after a node failure or in a full cluster. The autofission-runtime class is negative and also uses preemptionPolicy: Never; configure Fission to use it as shown under installation. Accurate resource requests remain essential because Autofission budgets requests like the scheduler rather than measuring live CPU or memory usage.

Treat permission to set the opt-in label as permission to consume the cluster's elastic budget. In a multi-tenant cluster, restrict that label with your admission policy. Autofission does not implement cross-Function quotas or fair sharing.

Operations

The chart runs one replica with a Recreate strategy, preventing rollout overlap within one installation without requiring leader election. After a node or cluster restart, the Deployment restores its controller replica and Autofission rebuilds its state from the API.

Useful checks for the default release name, namespace, and ServiceAccount:

kubectl rollout status deployment/autofission --namespace fission
kubectl auth can-i list pods \
  --as=system:serviceaccount:fission:autofission --all-namespaces
kubectl auth can-i patch functions.fission.io \
  --as=system:serviceaccount:fission:autofission --all-namespaces
kubectl auth can-i get secrets \
  --as=system:serviceaccount:fission:autofission --all-namespaces

With the default chart-managed RBAC and no additional bindings, the rollout should complete; the Pod-list and Function-patch checks should say yes, and the Secret check should say no.

Before uninstalling, remove the opt-in labels and set each Function's MaxScale to the limit you want to retain: stopping Autofission does not restore older values.

Uninstalling removes the controller resources but retains the runtime PriorityClass because Fission may still reference it during future cold starts. Fission CRDs, Functions, Environments, and namespaces remain untouched. Remove the PriorityClass manually only after ensuring that Fission's global runtimePodSpec and all Function or Environment PodSpecs no longer reference it.

Compatibility and limitations

  • Only Fission newdeploy Functions are managed. poolmgr, empty executor, and container are rejected as opt-in configuration errors.
  • Every managed Function receives the full capacity it could use by itself. This preserves burst capacity, but simultaneous cold bursts can temporarily create Pending Pods. Later cycles account for scheduled peers and reduce the limits; Autofission is not a fairness scheduler.
  • Fission and Kubernetes require a positive HPA maximum, so capacity below one replica produces MaxScale=1. An explicit MinScale is honored even when it exceeds currently free capacity.
  • The Fission v1 API is tested end to end with Fission 1.27.0. Environment resource inheritance follows Fission's override semantics and is covered by unit tests.
  • Capacity includes CPU, memory, and Pod slots. It does not model storage, GPUs and other extended resources, quotas, topology or affinity, per-Function scheduling constraints, image architecture, or unscheduled third-party Pods.
  • Tainted nodes are excluded unless --include-tainted-nodes is explicitly set. Only enable it when Fission runtime Pods actually tolerate those taints.
  • Extra sidecars and runtime PodSpec overhead are learned from a running Function Pod. Before the first replica, the estimate consists of the resolved runtime container plus configured fetcher requests.
  • Low, non-preempting priority prevents Function Pods from evicting existing workloads. It cannot prevent node-pressure eviction when requests are inaccurate or nodes run at their physical limit; reserve headroom and set accurate requests.
  • Very large clusters should benchmark API-server load and controller memory before shortening the default interval.
  • Cluster-scoped RBAC and PriorityClass names use the release name, so reusing a name in another namespace causes collisions. Prefer one Autofission installation per cluster. Multiple installations require unique fullnameOverride and priorityClasses.*.name values plus selectors that assign disjoint sets of Functions to their managedLabel and managedValue pairs; two controllers must never manage the same Function.

Troubleshooting

Autofission is NotReady — inspect controller logs. A 403 usually indicates missing custom RBAC or a ServiceAccount mismatch. A 409 means the Function changed after it was listed; the daemon will retry it on the next cycle. Other errors name the Function where possible.

The calculated limit is smaller than expected — check cordons, Ready status, taints, Pod requests, Pod slots, fetcher values, and per-node fragmentation. Capacity cannot combine spare CPU and spare memory located on different nodes.

Function Pods remain Pending — verify the runtime PriorityClass, taints/tolerations, node selectors, architecture, quota, storage, and concurrent managed Functions. Those constraints can be stricter than Autofission's current capacity model.

The HPA has not changed yet — Autofission patches the Function CR. Fission's executor reconciles that change into the HPA asynchronously; controller readiness confirms only that Autofission completed a full reconciliation successfully within the configured probe age.

Metadata

Release files for autofission 0.0.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for autofission 0.0.3
File Size Uploaded
autofission-0.0.3.tar.gz 66.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for autofission 0.0.3
File Interpreter ABI Platform
autofission-0.0.3-py3-none-any.whl Python 3 none any Details

Total release size: 91.5 kB

Release files / autofission-0.0.3.tar.gz

Download URL autofission-0.0.3.tar.gz
Size 66.0 kB
Tags Source
SHA-256 checksum
How to use checksums
4bea515ba489564953dd41a048035ae9d7050882a8e2abeb36fc33826b15e71a
BLAKE2b-256 checksum
How to use checksums
0e3e58bc0978d431ad5f0e68299c77750b1256bea32df87f38991d1cdcbdb8e5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 21, 2026.

Transparency log

Release files / autofission-0.0.3-py3-none-any.whl

Download URL autofission-0.0.3-py3-none-any.whl
Size 25.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2200341785a8d6912b516cb3dc015d4db6c490a96910ef0ba0af5a3583e881c7
BLAKE2b-256 checksum
How to use checksums
11c6ad6268373df28d1d085afec81f3a43a8456be655e310829a6993994487b5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 21, 2026.

Transparency log

Release history Release notifications | RSS feed

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

This release

0.0.3 This release

2 release files

0.0.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page