v2 Single-host Concurrency

The current Server can serve multiple users within one process, but deployment still uses one worker. Concurrency lets other tasks progress while one task waits for a model or storage. It does not automatically parallelize database writes or guarantee that more execution slots will improve throughput.

Commits and leases

A lease means a worker currently has authority to execute a Run; fencing checks reject writes from a worker whose authority has expired. Once a write passes the check and enters the protected section, the Scheduler preserves its authority until it finishes.

In-memory and file Schedulers hold the global execute_fenced lock briefly to check and update protection records. They do not block other Runs for the entire storage write.

  • Different Runs may commit concurrently; commits within one Run remain ordered.
  • A committing Run cannot be reclaimed, cancelled, or have its lease released; its tenant slot is not freed early.
  • Heartbeats may renew during writes; admitted protected writes retain their original lease authority.
  • Protection is removed after success, failure, or repeated cancellation. Shutdown waits for protected writes to exit.
  • Claim waiters wake at relevant lease expiry without depending on another request to notify them.

File Scheduler state remains atomically and sequentially persisted; SQL SessionStore writers still execute transactions sequentially. Reducing Scheduler-level serialization does not change each store’s guarantees.

Server configuration

Server keeps a long-lived Application. Chat and managed Applications have separate queues and a shared SchedulerQuotaGroup, so total quotas do not multiply with package count.

Environment variable Default Purpose
SAGE_SERVER_MAX_CONCURRENT_RUNS 8 Total executing root Runs
SAGE_SERVER_MAX_CONCURRENT_RUNS_PER_USER 2 Atomic per-user execution quota
SAGE_SERVER_MAX_PENDING_RUNS 1024 Shared pending-queue limit

This is an adjustable single-host acceptance starting point, not a capacity promise:

SAGE_SERVER_MAX_CONCURRENT_RUNS=32
SAGE_SERVER_MAX_CONCURRENT_RUNS_PER_USER=4
SAGE_SERVER_MAX_PENDING_RUNS=1024

Reduce execution concurrency when model rate limits or database capacity become bottlenecks, and observe queueing and P95/P99 latency. A longer queue does not increase throughput. Direct Builder users configure max_concurrent_runs, max_concurrent_runs_per_tenant, and max_pending_items under execution.scheduler. See resource management for Job limits.

Validation and deployment boundary

python -m pytest tests/sagents/v2/test_single_host_fencing.py
python scripts/v2/benchmark_v2_single_host.py --sessions 64 --writes 8 --io-ms 2
python scripts/v2/benchmark_v2_single_host.py --sessions 128 --writes 8 --io-ms 2

Tests cover parallel cross-Run commits, same-Run ordering, renewal, reclamation, quotas, cancellation, shutdown, and the Dispatcher → LeaseFencedSessionStore → SessionStoreCoordinator path. Benchmarks simulate slow asynchronous storage without real models or databases. Historical pass counts and synthetic latency are not production-capacity guarantees.

Server still requires one worker. Persistence, lease fencing, in-process event subscriptions, and JobRuntime are separate guarantees. Using MySQL does not establish multi-process or multi-host execution support.


Sage documentation for the current repository layout. Source available under the MIT license.

This site uses Just the Docs, a documentation theme for Jekyll.