Controls and verification
Independent controls
Section titled “Independent controls”| Control | What it measures | When to enforce |
|---|---|---|
| Cloud detection | Automated request evidence | After reviewing observed traffic; enable in SDK and dashboard |
| Local rules | Your application policy | After validating the rule against authenticated application state |
| Shared quota | Request admissions per account and optional session | When you need a request allowance across replicas |
| Concurrency | Simultaneously protected work | When all work, including streaming, is inside the lease lifecycle |
| Model budget | Reserved and confirmed tokens / configured cost | When every model attempt has conservative bounds and reliable accounting |
Shared controls use trusted server identities and stable secret-derived pseudonyms. Do not take an account ID or paid-plan claim directly from browser input. Keep property, rule IDs, secrets, and policy consistent across replicas. Policy changes may require a new rule ID and fresh counters; coordinate them deliberately.
Outages and cancellation
Section titled “Outages and cancellation”Cloud detection fails open by default: a detector outage lets the request proceed with degraded protection. Local rules and shared controls have separate failure policies; an unavailable cloud detector does not disable an enforced local rule.
Quota, concurrency, and budget controls default to observe/open. Enforced exhaustion denies work; unavailable shared state follows that control’s failure policy. Choosing closed can reject otherwise valid requests when the state service is unavailable. Choosing open prioritizes availability and cannot guarantee a hard cap.
For concurrency, confirmed completion releases capacity. Cancellation, errors, or lease loss can retain capacity until the configured maximum runtime because a disconnected client is not proof the provider stopped. Your handler must honor cancellation and bound upstream execution.
Model budgets and usage
Section titled “Model budgets and usage”Reserve a conservative maximum before each provider attempt, including retries, fallbacks, and tool-loop model calls. Supply a versioned price catalog and enforce input/output bounds in the application. SDK configuration alone does not constrain the provider.
Confirmed final usage reconciles the reservation. Missing usage, cancellation, or failed settlement retains the reserved maximum. Unknown usage is not zero usage. An overrun can be recorded but cannot reverse an already-billed model call.
The SDKs provide an Ollama final-usage normalizer. Other provider clients require application integration and verification of their final usage fields. Provider pricing, hidden retries, and token estimation affect the accuracy of any budget. These controls do not change your WebDecoy subscription or guarantee a provider invoice amount.
Data and reporting
Section titled “Data and reporting”Admission uses request metadata and local policy results. Prompts and model responses are not needed in detection or usage payloads. Shared identities are pseudonyms and remain correlatable; they are not anonymous. Keep raw application context and credentials out of custom reasons, route names, and logs.
Outcome and model-attempt reports are bounded, best-effort telemetry. An absent report does not prove a model call did not happen. Let completion hooks finish and drain reporting during graceful shutdown; reporting failures must not trigger another model call.
Verify your integration
Section titled “Verify your integration”Use a deterministic fixture or local model first, so validation does not consume paid model credits.
- Allowed request: authenticate, send a valid request, and confirm the handler runs once and the decision appears in the correct property’s AI Protection page.
- Denied request: trigger a known local policy denial and verify the model handler never starts. Check the reported application decision.
- Detector outage: simulate an unavailable detector in a test environment. Verify the configured fail-open behavior and degraded coverage.
- Shared quota: exhaust an enforced allowance across two application instances. Confirm 429 and retry guidance before inference. Admission recovery does not deduplicate model execution.
- Concurrency: hold a stream open and attempt another request above the limit. Verify completion releases capacity and disconnect handling follows the lease policy.
- Budget: exceed an enforced reservation limit and confirm no provider call starts. Separately test confirmed usage, unknown usage, and settlement failure.
- Dashboard: review allowed/denied decisions, degraded checks, model attempts, and known versus unknown usage. Confirm your application logs agree with the reported outcomes.
Test each configured outage policy separately. Do not infer budget or concurrency enforcement from a successful bot-detection test.