Polish
Polish is a declarative architecture language for people and LLMs. Describe
pages, services, data stores, and infrastructure in a .polishd file, check
their relationships, then run request scenarios against the resulting graph.
This repository contains version 0.7: a Python compiler and deterministic functional simulator. It does not provision infrastructure or send real network requests.
Problems Polish helps solve
Reuse an architecture when the use case changes
You have been vibe coding an ecommerce app with an LLM and refining its architecture along the way. Now you want to reuse that design for a ticketing site: keep the account and checkout flows, replace the catalog with events, and change the storage and access rules.
A .polishd file gives you and the LLM an explicit starting point. Copy the
specification, change the components and relationships, then rerun the scenarios.
The compiler checks the declared relationships; simulations show whether requests
still reach the intended services and data stores under the new assumptions.
Keep scenarios for behavior that must survive the change, such as preventing a
regular user from reaching an admin page.
Polish currently checks specifications you or an LLM write. It does not extract architecture from application code, generate the application, or verify that a deployment matches the specification.
Catch an infrastructure choice that conflicts with the design
You add path-based routing for /api/*, then discover that your load balancer is
an AWS Network Load Balancer. Polish rejects path rules on its NLB profile at
compile time:
load_balancer Ingress {
provider = aws
product = nlb
route "/api/*" to Apps
}
With Apps declared in the surrounding architecture, this produces E_VENDOR:
NLB cannot inspect HTTP paths; use routes_to. One solution is to change the
component to product = alb, configuring TLS termination if it accepts HTTPS.
Another is to keep the NLB forwarding traffic and put the path rules in an HTTP
router behind it. AWS documents the listener capabilities for
ALBs
and NLBs.
Test capacity assumptions before provisioning
Declare a workload and compare it with your capacity assumptions: request rate, CPU cost, instance counts, connection limits, database throughput, or queue consumer rate. Scenarios can expose a service that needs more instances, a queue whose backlog grows, or a data store whose declared request budget is exceeded. Change the design and rerun the same workload to compare outcomes. See scaling, resource budgets, and AWS processing for examples.
These are deterministic models of declared assumptions, not cloud benchmarks. A passing scenario means the modeled request satisfies the modeled constraints; it does not establish production performance or reliability.
Run it
Requires Python 3.11 or newer.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[dev]'
polish check examples/shop.polishd
polish simulate examples/shop.polishd
polish simulate examples/shop.polishd --scenario "Product page calls the API end to end"
polish simulate examples/shop.polishd --json
pytest -q
After installing, without activating the virtual environment, use .venv/bin/polish and .venv/bin/python -m pytest.
python -m polish is also supported.
If editable installation is unavailable in your environment, use a regular
installation (python -m pip install .). When editing the compiler, use .venv/bin/python -m polish from the project
root to run the current source, or reinstall with python -m pip install . to
refresh the installed polish command.
Commands return exit code 0 when checks or scenario expectations pass, and 1
for compilation errors, failed expectations, unreadable input, or no matching
scenarios. An expected denial or expected TLS error is a passing scenario.
--json emits a single JSON object with ok, diagnostics, and results.
Example
architecture CatalogSite {
artifact_store Assets { kind = object_store tls = true }
frontend Site {
rendering = static
loaded_from Assets
page Product
Product invokes Catalog.get_product
}
cdn Edge {
tls_termination = true
route "/api/*" to Balancer
}
load_balancer Balancer { routes_to Servers }
compute_pool Servers { instances = 2..10 }
service Catalog {
hosted_on Servers
operation get_product {
method = GET
path = "/api/products/*"
reads Products.Items
}
}
database Products {
kind = relational
read_replicas = 2
hosted_on Storage
table Items { id: uuid primary_key name: string }
}
database_cluster Storage { engine = postgres replicas = 2 }
}
scenario "Load product details" {
request {
entry = Site.Product
action = Catalog.get_product
via = Edge
protocol = https
method = GET
path = "/api/products/123"
}
expect {
outcome = success
reaches Catalog.get_product
accesses Products.Items
}
}
This scenario loads the page from its artifact store, checks that the page may
invoke the operation, and traverses the API route to the database access. Page
loading alone does not automatically invoke every service the page is allowed
to use. action selects an operation and via selects its network entry point.
See the ecommerce example for navigation, checkout, authorization, key-value storage, and nine executable scenarios. The invalid example demonstrates the compiler rejecting two loading sources.
Language rules
Each file contains one architecture Name { ... } followed by zero or more
scenario "Name" { ... } blocks. Identifiers are case-sensitive. Properties use
=, relationships use named verbs, and comments start with // or #.
Values include quoted strings, bare symbols, numbers, booleans, and inclusive
integer ranges such as 2..10. General variables, expressions, imports, and
user-defined types are not implemented yet.
Component kinds are built in. Nesting means containment: a page belongs to a frontend; it does not inherit from the frontend as a class. Pages inherit their frontend's rendering settings, loading source, and access restrictions.
| Component | Allowed children | Selected properties |
|---|---|---|
frontend |
page |
rendering = static or server |
service |
operation |
tls |
database |
table for relational; document or collection for document storage; keyspace for KV |
kind, read_replicas, partitions, durability and replication settings |
artifact_store |
none | kind = object_store or artifactory, tls |
cdn |
none | tls, tls_termination, upstream_protocol, request_collapsing, shielding, rate_limiting, bot_mitigation |
gateway |
none | tls, tls_termination, upstream_protocol, rate_limiting |
load_balancer |
none | tls, tls_termination, upstream_protocol |
cache |
none | tls, tls_termination, upstream_protocol, ttl |
compute_pool |
none | instances = 2..10, tls |
database_cluster |
none | engine, replicas, partitions, copies, failure_domains, storage, tls |
keyspace |
none | key_type, value_type |
role |
none | none |
An operation can declare method and path together to expose an HTTP endpoint.
Methods are GET, POST, PUT, PATCH, DELETE, HEAD, and OPTIONS. A service without
HTTP operations can serve as a page renderer or a logical service dependency.
Fields support uuid, string, decimal, integer, boolean, datetime, and
json, with optional primary_key metadata. No records or queries are executed.
Relationships
Product navigates_to Cart
Product invokes Catalog.get_product
Inside a component, omit the source to refer to that component:
service Catalog {
hosted_on Servers
operation get_product { reads Products.Items }
}
References resolve from the enclosing scope outward, then at the architecture
root. Use qualified names such as Site.Product or Catalog.get_product in
scenarios.
| Relationship | Source → target |
|---|---|
navigates_to |
page → page |
invokes |
frontend, page, service, operation → service, operation |
reads, writes |
service, operation → database, table, document, collection, keyspace, cache |
loaded_from |
frontend, page → artifact store for static rendering; service or load balancer for server rendering |
hosted_on |
service, self-hosted cache, self-hosted gateway → compute pool; database → database cluster |
routes_to |
CDN, gateway, load balancer, cache → network component, compute pool, service, operation, artifact store, database, database cluster |
Each frontend needs exactly one loading source. A page can override its inherited source with exactly one local source. Two local sources are an error, even if they point to the same component. Each service, database, and self-hosted cache needs exactly one compatible host. Managed cache profiles omit the host. Services cannot invoke frontends; returning a response is implicit and does not require an invocation edge.
route "/api/*" to Gateway is a path-specific routing edge. Patterns use shell
wildcards (*, ?, [abc]), with * also matching slashes. The route with the
longest literal prefix wins; equally specific matches fail as ambiguous.
routes_to Target is a fallback used when no explicit route matches.
At a compute pool, the simulator selects an operation from its hosted services using the request method and path. An exact endpoint match wins over wildcard matches; multiple remaining matches fail. Explicit routing to an operation selects that operation directly.
Access and TLS
page Checkout { requires authenticated }
page AdminPanel { requires role Admin }
Requirements also work on frontends, services, and operations. Parent
requirements apply to children. Role requirements need both
authenticated = true and the matching actor. Requests are anonymous by
default; specifying a role alone does not authenticate a request. All declared
requirements must be satisfied. Only one actor role is supported per request.
Network components accept HTTP by default. HTTPS requires tls = true or
tls_termination = true. A forwarding component that terminates TLS defaults to
HTTP upstream; upstream_protocol = https starts another TLS connection. Without
termination, HTTPS is preserved; changing it to HTTP causes
TLS_TERMINATION_REQUIRED. TLS support must be declared on every modeled
receiver of an HTTPS network hop.
Scenarios
A request can describe one of three flows:
- A network request:
entry = Edge,method = GET,path = "/api/products/123". - Navigation:
entry = Site.Landing,destination = Site.Checkout. - A frontend action:
entry = Site.Product,action = Catalog.get_product,via = Edge, plus the HTTP method and path.
protocol defaults to https, method to GET, and path to /. Navigation
finds an authorized path through declared links and loads the pages along it.
Actions and navigation are separate scenarios in this version.
Expectations require outcome = success, denied, or error. Optionally specify
an exact error code with error = TLS_UNSUPPORTED. Add reaches Component and
accesses Database.Table assertions to check the trace. Assertions also apply to
expected-error scenarios: they are not skipped on failure.
Service and database connections
Service and operation dependencies traverse a connection path before execution.
Use an optional block on invokes, reads, or writes:
invokes Inventory.reserve { via = InternalLB protocol = https }
reads Products.Items {
via = DataLB
protocol = tls
path = "/products"
read_from = replica
consistency = strong
}
writes Orders.Purchases { protocol = tls durability = replicated }
via starts the connection at a gateway, load balancer, or another permitted
network entry. Routing must reach the target workload or its declared host;
reaching a different host fails with WRONG_DESTINATION. Calls without via
connect directly to the target's host. The compiler still accepts existing bare
relationships: service calls default to http, database calls to tcp.
Service calls accept http or https; database calls accept tcp or tls
(encrypted TCP). Both the receiving host and workload must declare TLS support
when encryption reaches them. Forwarders can terminate TLS and optionally
re-encrypt upstream. Protocol families cannot change from HTTP to TCP or vice
versa. CDN and gateway nodes only support HTTP-family connections.
Connection path selects an infrastructure route. For service calls it defaults
to the target operation's declared path, otherwise /. For database connections
it is a logical routing label, not a SQL query or an actual TCP URL. Dependency
routing targets the named operation, rather than redispatching the caller's HTTP
method/path against the callee. A service target executes its service-level
requirements and dependencies; name an operation to execute that operation.
Each connection has its own routing-cycle check, so two services can legitimately
share a pool. Recursive execution still fails with REQUEST_CYCLE. Calls retain
the scenario identity. Service-level dependencies run before operation-level
ones, in declaration order, and an error stops subsequent execution. Data access
is only recorded after transport and read/write requirements pass.
See storage.polishd for a checkout that calls inventory through an internal load balancer, then uses all three database kinds. It also contains expected failures for missing TLS, inconsistent replica reads, and insufficient write durability:
python -m polish simulate examples/storage.polishd
Scale, storage, and durability
These properties define Polish contracts, not vendor-specific guarantees.
| Database kind | Data children | Partitioning |
|---|---|---|
relational |
table |
Optional count or managed |
document |
document, collection |
Optional count or managed |
key_value |
keyspace |
Optional count or managed |
A keyspace requires key_type (string, integer, uuid) and value_type
(string, integer, decimal, boolean, json, bytes). The compiler rejects
children incompatible with their database kind. A database may also be accessed
directly without declaring children.
database Sessions {
kind = key_value
partitions = managed
replication_factor = 3
read_replicas = 2
replication = asynchronous
durability = disk
hosted_on SessionStorage
keyspace Tokens { key_type = string value_type = json }
}
database_cluster SessionStorage {
storage = persistent
replicas = 2
copies = 3
partitions = managed
}
instancesaccepts a positive count or ordered positive range.partitionsaccepts a positive count ormanaged; omission means unspecified.replication_factorcounts all copies per partition, including the primary. It defaults to1 + read_replicas.read_replicasdefaults to zero and cannot exceed the total copies minus the primary. Copies need not all serve reads.- Host
replicasis read-replica capacity; hostcopiesis total copy capacity per partition. Explicit host counts bound the database's requested counts. Managed partition counts are opaque and are not compared numerically. - Database
failure_domainsdeclares the number of distinct placement domains. It cannot exceed the copy count or an explicitly declared host domain capacity. It does not promise survival under a particular failure; placement and quorum availability are not simulated. durability = memory(the default) promises no persistent write acknowledgement.diskrequiresstorage = persistenton the host.durability = replicatedrequires persistent host storage, at least two copies,replication = synchronous, andwrite_ack = quorumorall. Quorum meansfloor(copies / 2) + 1acknowledgements; all means every copy.primaryis the default acknowledgement policy. Quorum/all policies require synchronous replication and at least two copies even when durability is notreplicated.backups = truerequires a positivebackup_retention_days. Retention without enabled backups is an error. This validates a backup policy; it does not execute backup or restore jobs or guarantee an RPO/RTO.
A write may require a minimum durability using writes Store { durability = disk }.
The simulator checks the ordering memory < disk < replicated and traces the
number of configured acknowledgements. Copies alone do not imply persistent
storage, synchronous replication, or successful backup.
Reads default to read_from = primary and consistency = eventual. A replica read
requires at least one declared read replica. In this simplified contract,
consistency = strong on a replica requires synchronous replication; a primary
read is accepted. This models a declared capability, not a proof of database
linearizability, transactions, isolation, or actual replication lag.
Missing host capacity means unspecified, not zero. Checks apply per database; shared-host resource reservations and aggregate capacity are not modeled yet.
AWS CPU capacity scenarios
Run .venv/bin/python -m polish simulate examples/scaling.polishd to see both
service and PostgreSQL overloads, with explicit arithmetic in the trace.
Add --json for structured per-resource capacity results.
The initial AWS catalog in polish/config/vendors/aws/ec2.json supports EC2 m7i.large,
m7i.xlarge, and m7i.2xlarge. vCPU and memory figures are sourced from
AWS's instance specifications.
These EC2 profiles support services and self-managed PostgreSQL. RDS uses separate
db.* profiles described below.
There is no AWS account access, deployment, pricing, or region-availability lookup.
Memory specifications also support the optional service memory model below.
compute_pool AppPool {
provider = aws
instance_type = "m7i.large"
instances = 1..4
cpu_utilization = 0.7
}
Declare cpu_ms = 10 on an operation, or on its service as a fallback. An operation
cost replaces its service fallback; it does not add to it. This represents total
vCPU-milliseconds per invocation, not elapsed request latency. Explicit zero is
allowed; an omitted cost is not silently assumed to be zero.
In the scenario request, set rate_rps = 200 and optionally scale = minimum or
maximum. The default is minimum: a 1..4 pool uses one running instance.
Maximum evaluates the same workload with four instances, without simulating
autoscaling decisions or startup delays. Fixed instance counts are unchanged.
demand = requests_per_second × CPU_ms_per_request
capacity = running_instances × vCPUs_per_instance × 1000 × cpu_utilization
required_instances = ceil(demand / capacity_per_instance)
For the example: 200 × 10 = 2000 CPU-ms/s, while one 2-vCPU instance at a 70%
budget provides 1 × 2 × 1000 × 0.7 = 1400 CPU-ms/s. One fails; two suffice under
the model; four pass. cpu_utilization defaults to 0.7, a configurable modeling
assumption rather than an AWS guarantee. CPU budgets apply per pool.
Every executed operation and service dependency is charged once per execution per root request. Costs aggregate across services sharing a compute pool and across multiple reads/writes using the same database role. Scenarios run in isolation; multiple traffic classes, probabilistic branches, retries, variable fan-out, and background traffic are not combined automatically.
For PostgreSQL, declare an AWS profile on the database_cluster, use
engine = postgres, and put cpu_ms on each reads/writes relationship:
reads Products.Items { read_from = replica cpu_ms = 20 }
writes Products.Items { cpu_ms = 20 }
Replica reads share read_replicas running instances. The host's replicas
property is a capacity ceiling, not a count of active readers. Primary reads and
writes share exactly one primary. Increasing replicas can fix read overload but
cannot fix primary overload; use a larger instance or reduce workload/cost.
Primary required_instances is a CPU-equivalent count only, accompanied by
horizontal_scaling = false. This version requires a dedicated host cluster per
database for workload simulation; DB partitioning does not multiply capacity.
Without rate_rps, existing scenarios remain functional checks. With it, every
executed service/database workload needs an AWS host profile and explicit CPU
cost; incomplete models fail with SCALING_MODEL_INCOMPLETE. Database CPU scaling
currently supports PostgreSQL relational stores only. Static-only traffic is not
modeled. Overload returns CAPACITY_EXCEEDED; expected overloads can be passing
tests. Capacity is evaluated after functional traversal, before returning success.
The calculation assumes perfectly balanced, parallelizable CPU work and constant
costs. Measure those costs for your application and chosen machine profile.
vCPU count alone is not a performance benchmark. Disk IOPS, locks, network bandwidth, replication overhead, and tail latency are not
calculated; memory and connection budgets are optional models described below; include known CPU overhead in the declared cost. This is a
steady-state feasibility model, not a throughput guarantee or a latency simulator.
Arithmetic formulas are implemented by the simulator; .polishd currently supplies
their numeric inputs rather than supporting general expressions.
Memory and database connection budgets
Run python -m polish simulate examples/budgets.polishd for five scenarios covering
CPU overload, connection-pool oversubscription, memory overload, and scale-out.
JSON output includes a budgets array alongside CPU capacity records.
These checks opt in through declarations and run on scenarios with rate_rps.
Existing CPU-only and functional scenarios retain their previous behavior.
service Checkout {
hosted_on Apps
base_memory_mib = 512
connection_pool_size = 20
operation buy {
method = POST path = "/buy"
cpu_ms = 10
duration_ms = 100
memory_per_request_mib = 2
writes Products.Items { cpu_ms = 0.1 }
}
}
Memory uses the following steady-state estimates:
concurrent requests = rate_rps × duration_ms / 1000
inflight memory = concurrent requests × memory_per_request_mib
base memory = instances × sum(resident services' base_memory_mib)
available memory = instances × instance_memory_GiB × 1024 × memory_utilization
Set memory_utilization on the compute pool (default 0.8 when the model is enabled)
to reserve headroom for the operating system and other overhead. The model counts
base memory for every resident service, including services not called in the
scenario, assuming one process of every resident service on every pool instance.
Every resident must declare base_memory_mib when memory modeling is enabled;
place unsupported resident types such as caches in separate pools.
Operations declare duration_ms and memory_per_request_mib, or inherit those
values from their service. Duration is elapsed time, separate from cpu_ms.
Explicit zeros are allowed. Inflight costs sum across executed operations; each
invocation's memory footprint must also fit a single instance alongside resident
base memory. Missing required inputs fail explicitly. The model assumes balanced
traffic and fixed costs, and does not predict peak concurrency, garbage collection,
allocator fragmentation, leaks, or PostgreSQL's internal memory usage.
For connections, set database max_connections and optional
reserved_connections (default zero). Set each calling service's
connection_pool_size, interpreted as a maximum per application instance per
database role. A role is the primary or the read-replica group.
potential connections = sum(calling service instances × connection_pool_size)
available per DB instance = max_connections − reserved_connections
Repeated reads/writes from the same service to the same role reuse the same pool
budget. Different services add to the budget. Replica pools are assumed perfectly
balanced across readers: each reader needs ceil(total / read_replicas) slots.
The primary has one budget shared by primary reads and writes. Only services whose
database edges execute in the scenario contribute pools; background or idle pools
outside that path must be accounted for through reserved connections or separate
scenarios. This does not model PgBouncer or multiplexing.
The example has 200 database slots, with 20 reserved. Four application instances
with pools of 20 need at most 80 slots. Twenty instances need 400 slots and fail
with CONNECTION_BUDGET_EXCEEDED, even though their CPU budget now suffices.
Reducing the pool size to 9 brings that budget to 180 and passes the model.
This is a conservative potential oversubscription check, not observed runtime connection exhaustion: it assumes configured pools could fill. Actual open connections, checkout wait time, and connection-hold duration are not simulated. CPU feasibility is evaluated first, followed by memory and connection budgets; a CPU failure can therefore stop the scenario before budget results are emitted.
Cloud vendor catalogs and AWS products
Vendor-specific definitions live under their own folders, separate from the core language definitions:
polish/config/
vendors.json
vendors/
aws/
ec2.json # EC2 instance hardware
rds.json # RDS PostgreSQL and db.* instance classes
dynamodb.json # Managed storage and capacity-unit rules
elb.json # ALB/NLB protocol and routing capabilities
vendors.json lists the product files to load for each vendor. Custom config
directories must include that manifest and its referenced files. The old root
aws.json has been removed; migrate custom copies to the new layout. Catalogs
ship with the Python package. AWS, self-hosted software, Kong, and Fastly profiles
are implemented; adding another folder alone does not implement its behavior. Instance sizes, product properties,
defaults, restrictions, and capacity-unit constants are data; execution remains
in the simulator. Product profiles apply when both provider = aws and a matching
product are declared.
Run the complete example:
python -m polish simulate examples/aws-products.polishd
RDS PostgreSQL
database_cluster Rds {
provider = aws
product = rds_postgres
instance_type = "db.m7i.large"
multi_az = true
}
Supported classes initially include db.m7i.large and db.m7i.xlarge, sourced from
AWS RDS hardware specifications.
EC2 names cannot be substituted for RDS classes, and RDS classes cannot host
compute pools. Engine defaults to postgres, storage to persistent, and TLS
support to true. Hosted relational databases default to TLS support and disk
durability. Explicit TLS settings still determine whether a modeled request passes.
Existing PostgreSQL CPU and connection-budget checks apply.
multi_az represents the Multi-AZ DB instance deployment, not an Aurora or
Multi-AZ DB cluster. Its standby adds no modeled read or write capacity. AWS
documents this distinction in its
Multi-AZ DB instance guide.
Declare read_replicas separately on the database. These default to asynchronous
replication; this profile rejects synchronous read-replica declarations. Failover,
standby replication cost, availability guarantees, engine versions, region support,
and service quotas are not simulated.
DynamoDB
database Sessions {
kind = key_value
partitions = managed
hosted_on DynamoTable
keyspace Tokens { key_type = string value_type = json }
}
database_cluster DynamoTable {
provider = aws
product = dynamodb
capacity_mode = provisioned
read_capacity_units = 100
write_capacity_units = 50
}
Here database_cluster represents a managed product endpoint, not a machine you
provision. One such host and its logical database represent one DynamoDB table's
capacity budget. Nested keyspaces/collections are schema views within that table,
not independently provisioned AWS tables. Use separate database/host pairs for
separate DynamoDB tables. Both key-value and document database kinds are supported.
Partitions are managed; instance types, CPU budgets, connection limits, and
user-managed replicas are rejected. Transport defaults to HTTPS. This version
models managed persistence at Polish's disk durability level; it does not model
DynamoDB replication internals or global tables.
reads Sessions.Tokens { item_bytes = 4096 consistency = strong }
writes Sessions.Tokens { item_bytes = 1500 }
Each edge models one nontransactional item operation per root request. Supply the
complete item size in bytes, including attribute-name overhead. Size must be
between 1 and 409600 bytes. For a scenario with rate_rps:
strong read units/sec = rate_rps × ceil(item_bytes / 4096)
eventual read units/sec = strong_read_units × 0.5
write units/sec = rate_rps × ceil(item_bytes / 1024)
These constants follow AWS DynamoDB capacity rules.
Thus 30 writes/sec of 1500-byte items consume 60 write units/sec and exceed a
50-unit provisioned budget. Demand aggregates across repeated accesses to the
same database; read and write budgets are separate. item_bytes is required for
workload simulation; cpu_ms and read_from are not valid on DynamoDB edges.
For capacity_mode = on_demand, omit provisioned capacities and optionally declare
max_read_units and max_write_units. A rate-based scenario requires an explicit
maximum for whichever operation it exercises. This bounds the simulation; it does
not assume immediate or unlimited scaling. Burst capacity, ramp-up, hot partitions,
indexes, transactions, scans, conditional-write behavior, retries, and billing are
not modeled. Throttling is a deterministic aggregate-budget failure, not a precise
prediction of which individual AWS requests would throttle.
AWS load balancers
load_balancer PublicALB {
provider = aws
product = alb
tls_termination = true
route "/api/*" to Apps
}
load_balancer InternalNLB {
provider = aws
product = nlb
tls = true
routes_to Rds
}
ALB accepts modeled HTTP/HTTPS traffic and path routes. HTTPS must terminate at
the ALB. NLB rejects path routes and forwards using routes_to; TCP/TLS and HTTP
payloads over those transports are supported by this simplified model. tls = true
without termination represents encrypted passthrough. It does not mean NLB
inspects HTTP paths. UDP, QUIC, listener ports, target groups, health checks,
certificates, LCUs, and load-balancer capacity limits are not modeled yet. See
AWS ALB listeners
and AWS NLB listeners.
No AWS resources are created or queried while compiling or simulating.
Caches, GraphQL, gateways, databases, and AWS processing
Both validation and executable models are provided for the products below. Models are deliberately explicit: declared throughput limits are workload measurements or planning assumptions, never invented vendor benchmarks.
| Component and product | Deployment | Executable checks |
|---|---|---|
cache, redis / memcached |
provider = self_hosted, required compute host |
Hit/miss paths, fallback calls, dataset size, operations/sec |
cache, elasticache_redis / elasticache_memcached |
provider = aws, no compute host |
Same cache model with managed hosting |
database_cluster, cockroachdb |
provider = aws or self_hosted |
Relational schema, transport, declared cluster operations/sec |
database_cluster, mongodb |
provider = aws or self_hosted |
Document schema, transport, declared cluster operations/sec |
gateway, graphql |
provider = self_hosted, required compute host |
HTTP routing, declared query-depth/cost limits |
gateway, gateway |
provider = kong, required compute host |
HTTP routing, request-rate and burst budgets |
gateway, api_gateway |
provider = aws |
HTTP routing, request-rate and burst budgets |
cdn, cdn |
provider = fastly |
Edge routing, request-rate and burst budgets |
gateway, signal_sciences or next_gen_waf |
provider = fastly |
Inline policy routing, request-rate and burst budgets |
stream, msk |
provider = aws |
Broker/replication constraints, message throughput, backlog growth |
queue, sqs |
provider = aws |
Standard/FIFO configuration, message throughput, backlog growth |
batch_cluster, emr |
provider = aws |
Spark/Hadoop configuration, worker CPU capacity |
AWS CockroachDB/MongoDB profiles describe software deployed on AWS, not AWS-native managed database products. They do not imply Atlas or CockroachDB Cloud support. Signal Sciences is exposed under Fastly, consistent with Fastly's Next-Gen WAF documentation. The WAF node abstracts an inline policy check; attack detection and vendor agent deployment are not reproduced. GraphQL is an HTTP layer, not a cloud vendor.
python -m polish simulate examples/platforms.polishd
python -m polish simulate examples/aws-processing.polishd
Caches and self-hosted hardware
compute_pool Local {
provider = self_hosted
vcpus = 4
memory_gib = 16
instances = 2
}
cache Redis {
provider = self_hosted
product = redis
hosted_on Local
capacity_mib = 512
dataset_mib = 100
max_ops_rps = 1000
invokes API.backend
}
Services access caches with reads Redis or writes Redis. Use scenario
cache_result = hit to serve a hit, or miss (the default) to execute the cache's
invokes fallback. A miss without a fallback fails explicitly. Hits skip fallback
CPU and database work; repeated cache accesses add operations to the cache budget.
Writes record access but do not mutate persistent state across scenarios. Cache
edges accept transport settings, not database consistency or durability options.
capacity_mib is the usable logical cache allocation, not raw instance RAM.
dataset_mib defaults to zero (an undeclared dataset), and exceeding the allocation
fails before fallback. A rate-based scenario requires max_ops_rps. These are
aggregate cache budgets, not inferred Redis/Memcached performance. Eviction,
TTL expiry, serialized keys, warming, cluster hashing, persistence, and inferred
hit ratios are not modeled. The selected hit/miss applies to all cache reads in
that scenario. Redis and Memcached share this intentionally generic cache model.
For ElastiCache, select the AWS product and omit hosted_on; capacity is still an
explicit assumption. See AWS's engine comparison.
Self-hosted compute pools use explicit vcpus and memory_gib, not an AWS instance
name. Existing service CPU and memory math applies. Database and cache declared
capacity budgets remain separate; the model does not infer their throughput from
CPU specifications or aggregate cache allocation against colocated service RAM.
MongoDB and CockroachDB
database_cluster Documents {
provider = self_hosted
product = mongodb
nodes = 3
max_ops_rps = 100
}
Attach a document database to MongoDB, or a relational database to CockroachDB.
nodes validates deployment size. max_ops_rps is a measured whole-cluster
budget; increasing nodes does not automatically multiply it. Each executed read
or write adds the scenario rate to the host's demand, including accesses from
different logical databases. Exceeding the budget fails. Wire-level query plans,
distributed transactions, MongoDB write concern, CockroachDB range placement,
replication protocols, elections, and hot shards are not simulated. Generic Polish
storage declarations do not prove those vendor-specific semantics. Their concepts
are described in CockroachDB deployment documentation
and MongoDB replication documentation.
Gateways, GraphQL, and rate limits
Set limit_rps for steady-state offered load and burst_limit for an instantaneous
burst. Scenarios supply rate_rps and optional burst_requests. Repeated traversal
of a policy node adds offered load. Exceeding either limit returns RATE_LIMITED
with outcome denied before the downstream request runs. At the limit, traffic
passes. No requests are partially admitted: an overloaded workload fails as a whole.
For GraphQL, set max_query_depth and/or max_query_cost; the scenario must supply
the corresponding query_depth/query_cost. Excess returns
GRAPHQL_LIMIT_EXCEEDED. The layer forwards to a service operation; that operation
declares resolver-like dependencies using invokes and data access edges. Polish
does not parse GraphQL documents, infer cost, validate a GraphQL schema, implement
federation, or infer N+1 fan-out. Gateway host requirements validate topology; the
gateway's own CPU overhead is not automatically charged to its host.
These common policy checks are not exact replicas of vendor algorithms. In particular AWS API Gateway uses token buckets and best-effort throttling, while Kong offers multiple rate-limiting plugins. This iteration checks steady-rate and burst budgets independently; token refill, time windows, distributed counters, per-identity quotas, and vendor licensing or plan availability are outside the model.
MSK, SQS, and EMR workload math
Services and operations may publishes_to / consumes_from a stream or queue,
or submits_to a batch cluster. Each edge represents one message/job per root
request. SQS/EMR use modeled HTTPS connections; MSK uses TLS. Missing throughput
inputs in a rate-based scenario fail rather than imply infinite capacity.
stream Events {
provider = aws product = msk
brokers = 3 partitions = 6 replication_factor = 3
max_messages_rps = 100
consumer_rate_rps = 80
max_backlog = 1000
}
queue Jobs {
provider = aws product = sqs queue_type = standard
max_messages_rps = 100
consumer_rate_rps = 5
max_backlog = 20
}
batch_cluster Analytics {
provider = aws product = emr framework = spark
workers = 2 worker_vcpus = 2 cpu_utilization = 0.7
}
MSK replication cannot exceed broker count. MSK/SQS max_messages_rps is a declared
aggregate throughput budget applied separately to publishing and consuming; it
is not an AWS service quota. consumer_rate_rps declares background drain; explicit
consumes_from edges add modeled consumption demand. Effective drain is capped at
max_messages_rps. Starting with an empty backlog:
backlog = max(0, publish_rate − effective_consume_rate) × window_seconds
The scenario's window_seconds defaults to 1. max_backlog is optional; if supplied,
exceeding it fails. A consume-only scenario tests capacity without requiring stored
messages. Partitions do not automatically multiply cluster throughput. Visibility
timeout, retention, queue type, and replication count are configuration metadata;
delivery order, duplicates, retries, actual consumer groups, retention expiry, and
Kafka offsets are not simulated. SQS standard queues can deliver duplicates and
out-of-order messages; see AWS SQS documentation.
An EMR submission supplies submits_to Analytics { job_cpu_ms = 1000 }:
batch demand = rate_rps × job_cpu_ms
batch capacity = workers × worker_vcpus × 1000 × cpu_utilization
At 3 jobs/sec, that is 3000 CPU-ms/sec. Two 2-vCPU workers at 70% offer 2800, so the scenario fails; three workers pass. Default utilization is 0.7. This assumes perfect parallelism and excludes driver overhead, Spark stages, shuffle, storage, memory, startup, and job scheduling. It is a workload feasibility model for an EMR-style processing cluster, not execution of a Spark/Hadoop job. All scenario state and budgets reset between runs.
Boundaries of this version
The simulator proves declared functional paths under the model's assumptions. It does not prove production availability, security, or performance. In particular:
- CPU capacity is modeled for explicitly declared workloads. Autoscaling timing, latency and physical failures are not simulated; queue backlog uses the aggregate model below. Storage configuration capacity is checked per database, not aggregated across tenants.
- CDN protection features, cache TTL, database engines, schema fields, and primary keys are typed or recorded metadata. Bot detection, caching, request collapsing, query semantics and version compatibility are not simulated. Explicit rate-policy budgets and cache hit/miss paths are simulated as described below.
- TLS checks cover request routing, page loading, service calls, and database connections. Certificates, TLS versions, real wire handshakes, and network ACLs are not modeled.
- Responses return implicitly. HTTP status codes, payloads, retries, queues, sessions, transactions, conditional branches, and reverse response paths are not modeled.
- Server rendering through a compute pool requires exactly one renderer service in that pool. Hybrid rendering can override settings per page, but streaming and hydration behavior are not modeled.
- There is no deployment backend, LLM integration, custom class inheritance, or persistent compiled artifact yet. JSON diagnostics provide an integration surface for an LLM to inspect and repair a specification.
Implementation
Language definitions live in polish/config/ and ship with the installed package:
| File | Definitions |
|---|---|
databases.json |
Database kinds and allowed children, database properties, storage enums, positive integer settings |
services.json |
Service and operation properties and containment |
hosts.json |
Host properties and permitted hosting pairs |
frontend.json |
Frontend/page properties, rendering sources, artifact store kinds |
network.json |
CDN, gateway, load balancer, and cache properties |
relationships.json |
Allowed source/target types, connection options and protocols |
language.json |
Boolean properties, field types, HTTP methods, shared enums and integer settings |
vendors.json |
Vendor catalog manifest |
vendors/aws/*.json |
AWS EC2, RDS, DynamoDB, and ELB profiles, capabilities, and source metadata |
errors.json |
Compiler, storage, simulator, and CLI error codes, message templates, and failure outcomes |
For example, databases.json declares the storage kinds and their child types:
"database_kinds": {
"relational": ["table"],
"document": ["document", "collection"],
"key_value": ["keyspace"]
}
The compiler loads definitions on each compilation. The grammar's component and relationship keywords come from these files too. To use a separate configuration, copy the whole directory and pass it explicitly:
cp -R polish/config /tmp/my-polish-config
python -m polish check examples/shop.polishd --config-dir /tmp/my-polish-config
python -m polish simulate examples/storage.polishd --config-dir /tmp/my-polish-config
The Python API accepts compile_source(source, config_dir=...). Custom directories
replace the bundled configuration completely; there is no implicit working-directory
override. Missing files and malformed definitions produce E_CONFIG diagnostics.
Run python -m polish from the repository to use edited source configuration, or
reinstall the package to refresh the installed CLI's bundled copy.
Configuration defines the vocabulary and permitted declarations. Execution algorithms—TLS handshakes in the model, routing, quorum calculations, durability requirements, and replica consistency checks—remain Python implementations. Adding a metadata property or database kind does not invent new runtime behavior. New executable component types, protocols, or durability semantics still require corresponding simulator support; changing an enum alone does not implement them.
Error definitions use stable rule identifiers. For example:
"TLS_UNSUPPORTED": {
"code": "TLS_UNSUPPORTED",
"phase": "simulation",
"message": "'{node_name}' does not accept TLS.",
"parameters": ["node_name"],
"outcome": "error"
}
Edit code or message to customize diagnostics. Simulation outcome may be
error or denied; failures cannot be changed into successes. Rule identifiers
with suffixes such as _2 distinguish different checks that share an emitted
error code. Keep identifiers, parameter contracts, and phases unchanged. Templates
support named placeholders and !r/!s, without attribute access, indexing,
format specifications, or executable expressions. A message may omit parameters.
The simulator retains the catalog from compilation, including custom
--config-dir definitions. If you change an emitted error code, update any
scenario expect { error = ... } assertions that reference it. Existing scenario
definitions remain in .polishd files. The conditions that detect errors still
live in the compiler/simulator; the catalog describes their diagnostics, not
arbitrary executable checks. E_CONFIG is a built-in fallback so a broken or
missing error catalog can itself be reported.
polish/grammar.lark: formal syntax, parsed using Lark.polish/model.py: typed component graph, fields, diagnostics, scenarios.polish/compiler.py: parsing, name resolution, type and constraint validation.polish/simulator.py: request traversal and scenario assertions.polish/storage.py: database contracts and connection-option validation.polish/cli.py: terminal and JSON output.tests/test_polish.py: positive and negative integration tests.
CloudFront, Cloudflare, S3, and ECR
See delivery.polishd for ten executable success and
failure scenarios. Profiles live in config/vendors/aws/delivery.json and
config/vendors/cloudflare/edge.json inside the polish package.
artifact_store Assets {
provider = aws product = s3
max_requests_rps = 100
}
cdn Edge {
provider = aws product = cloudfront
limit_rps = 200
routes_to Assets
}
frontend Shop { rendering = static loaded_from Edge }
Use provider = cloudflare product = cdn for Cloudflare. Both profiles default
to TLS termination with HTTPS upstream and support routing, TLS checks,
limit_rps, and burst_limit. These limits are explicitly declared scenario
policies, not vendor quotas. Every simulated request reaches the origin;
CDN cache hits, shielding, geographic latency, and cache eviction are not modeled.
S3 defaults to a REST endpoint. With endpoint = website, TLS defaults to false,
HTTPS fails, and only GET/HEAD requests succeed. Configure a CDN's
upstream_protocol = http to reach that endpoint. See the
AWS endpoint comparison.
A rate scenario reaching S3 requires max_requests_rps:
origin demand = scenario rate_rps × visits to that origin
failure when origin demand > declared max_requests_rps
This is an aggregate request budget, not an estimate of S3's actual performance. Object existence, bucket permissions, CloudFront OAC, and S3 durability are not simulated in this iteration.
ECR
is a container image registry, so use a service's image_from relationship:
artifact_store Images { provider = aws product = ecr max_pulls_rps = 5 }
service Api {
hosted_on Apps
image_from Images { authorized = true }
operation get { method = GET path = "/" cpu_ms = 1 }
}
A service may have one image source. Registries cannot be used as frontend
origins. The simulator checks HTTPS and workload authorization once per executed
service per scenario. authorized defaults to false and represents the declared
pull permission; it is independent of the request user's authentication.
An optional scenario image_pulls_rps models startup pull traffic, independently
of HTTP rate_rps. It requires max_pulls_rps; demands from distinct executed
services using the same registry are summed. These checks represent a deployment
precondition, not a fresh pull per application request. IAM evaluation, image
layers/tags, network transfer time, and registry storage capacity are not modeled.
License
Polish is licensed under the MIT License.
Metadata
Release files for polish-lang 0.7.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| polish_lang-0.7.0.tar.gz | 194.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| polish_lang-0.7.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 257.7 kB
Release files / polish_lang-0.7.0.tar.gz
| Download URL | polish_lang-0.7.0.tar.gz |
|---|---|
| Size | 194.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
87baf8d1d4c09e5ad88973148ec44855fe3d9447c60b315655efe4cdeb959cab
|
|
BLAKE2b-256 checksum How to use checksums |
139566648e6bd00cb9683699e281050db7cd6b2c99387c48a54919ac900aed25
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.3
|
Release files / polish_lang-0.7.0-py3-none-any.whl
| Download URL | polish_lang-0.7.0-py3-none-any.whl |
|---|---|
| Size | 63.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
9a4efa50d71bf49bb2bf9cee2ae555a201b93032eb6c13f2059b3ace685aba54
|
|
BLAKE2b-256 checksum How to use checksums |
5e05ada889b4eac037810b62b893d915bda0fc32860a5b9be84ce5f46d6280f7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.3
|