| Cost component | Monthly cost | Billing basis | Copy |
|---|---|---|---|
| {{ row.label }} | {{ row.value }} | {{ row.detail }} |
{{ item.title }}
{{ item.body }}
| Cost component | Monthly cost | Billing basis | Copy |
|---|---|---|---|
| {{ row.label }} | {{ row.value }} | {{ row.detail }} |
{{ item.body }}
A Lambda Functions bill is assembled from several meters rather than one flat price. Requests are counted separately from execution duration, while configured memory converts running time into gigabyte-seconds. Workloads that use Provisioned Concurrency, extra ephemeral storage, asynchronous payloads, or streamed responses can activate additional charges.
This matters most when a workload changes shape. Doubling invocations usually doubles request and execution quantities, but increasing memory can also shorten runtime. A provisioned window adds capacity cost even when traffic is quiet, and account-wide pricing tiers or allowances may already be partly consumed by other functions.
| Planning quantity | What changes it | Common source |
|---|---|---|
| Request units | Invocation count and asynchronous event size | Invocation metrics plus retry and payload forecasts |
| On-demand GB-seconds | Invocations, billed duration, and configured memory | Lambda reports or a measured load test |
| Provisioned charges | Traffic share, duration, configured concurrency, and enabled hours | Alias or version configuration and traffic schedule |
| Optional meters | Storage above 512 MB and streamed bytes above 6 MB per response | Function configuration and response measurements |
A useful estimate therefore begins with the billed quantities, not only the nominal request count. Configured memory must replace peak memory use, duration should be the billed duration for the selected architecture, and any no-charge allowance must be limited to the amount still available to this workload.
The result covers the modeled Lambda Functions meters only. Logs, API Gateway, queues, databases, data transfer, NAT, support, taxes, credits, private pricing, and newer Lambda products or features need separate estimates when they apply.
Start with a representative workload, then replace every assumption that can be measured or quoted for the target Region.
The monthly total is a planning estimate for the active meters. Read it together with the ledger basis: a low request charge can sit beside a much larger duration or provisioned-capacity charge, and a zero line may reflect an entered allowance rather than zero usage.
Execution duration is priced from allocated memory and billed time. On-demand and provisioned traffic are separated because they can use different duration rates, while Provisioned Concurrency capacity depends on configured environments and enabled time rather than invocation count.
The monthly total is the sum of request, execution, provisioned-capacity, storage, and streaming charges after the entered allowances are applied.
| Symbol | Meaning | Unit |
|---|---|---|
| No, Np | On-demand and provisioned-window invocations | count |
| to, tp | Average billed duration for each traffic group | ms |
| M | Configured memory allocation | MB |
| C | Provisioned concurrency | environments |
| H | Provisioned enabled time | hours |
| Go, Gp, Gc | On-demand duration, provisioned duration, and configured capacity | GB-s |
Each cost component multiplies its billable quantity by the active rate. Entered request and on-demand compute allowances are subtracted before their rates are applied. Provisioned duration and configured capacity do not use the entered on-demand compute allowance.
The built-in card uses these documented US East example rates. A custom rate card replaces them when enabled.
| Meter | x86 | Arm | Rate unit |
|---|---|---|---|
| Requests | $0.20 | $0.20 | per million request units |
| On-demand duration | $0.0000166667 | $0.0000133334 | per GB-s |
| Provisioned duration | $0.0000097222 | $0.0000077778 | per GB-s |
| Provisioned capacity | $0.0000041667 | $0.0000033334 | per GB-s |
| Extra ephemeral storage | $0.0000000309 | $0.0000000309 | per GB-s |
| Response streaming | $0.008 | $0.008 | per GB |
Several meters use step or floor rules that are easy to miss in a simple request-times-duration estimate.
| Meter | Modeled rule | Boundary |
|---|---|---|
| Asynchronous requests | One request unit through 256 KB, then one additional unit for each started 64 KB chunk | At 256 KB the factor is 1; above 256 KB the next partial chunk counts |
| Ephemeral storage | Execution seconds multiplied by configured storage above 512 MB | Exactly 512 MB adds no storage quantity |
| Response streaming | Streamed invocations multiplied by bytes above 6 MB per response | Responses at or below 6 MB add no streaming overage |
| Budget | Monthly total minus the entered budget | A positive variance is over budget; zero or a negative variance is not over |
Intermediate quantities retain full numeric precision. Currency is rounded for display, so displayed line items may differ by a few cents from a total calculated from unrounded values.
Built-in rates are documented US East examples, not a live quote. Pricing can differ by Region, architecture, usage tier, account agreement, and date.
Three million x86 invocations at 120 ms and 1,536 MB produce 540,000 GB-s. With 400,000 GB-s and one million request units still available, the built-in example rates give about $2.33 for compute plus $0.40 for requests, or approximately $2.73 for the modeled month.
An average asynchronous event of 257 KB uses two request units because the first byte above 256 KB starts a 64 KB chunk. At 25% async traffic, only that quarter of invocations receives the 2× factor; compute invocations do not double.