Skip to navigation
Plugin Building MasterclassEnd-to-End Course

Test and Operate

Validate each module, the full execution path, retrieval, permissions, security boundaries, and production behavior.
View as Markdown

Test the plugin as a stack of contracts. A successful happy-path conversation does not prove that its authentication, mappings, policies, failure paths, and launch access are correct.

Test in Layers

1

Validate downstream access

Confirm the test identity, scopes, roles, tenant, environment, sample data, token expiration, and API limits.

2

Validate the connector

Test authentication with a safe read operation. Verify that the connection cannot reach unintended environments or records.

3

Validate each action

Test every input, request field, response schema, error response, and idempotency behavior in isolation.

4

Validate mappings

Confirm types, literal quoting, missing fields, nulls, empty arrays, output replacement, and field pruning.

5

Validate slots and resolvers

Test clear, ambiguous, invalid, omitted, and adversarial user phrasing. Confirm deterministic validation and recollection.

6

Validate compound execution

Test each control-flow branch, error handler, failure, partial completion, return mapping, and absence of leaked intermediate data. If you implemented a manual retry pattern, test its attempt limit and idempotency explicitly.

7

Validate the conversation process

Test collection order, confirmation, branching, content activities, context size, and final presentation.

8

Validate retrieval or system triggering

Test positive and negative utterances, webhook verification and payloads, or schedule timing, timezone, publication, and Active state.

9

Validate access end to end

For conversational plugins, confirm launch configuration and dependency Use access. For system triggers, confirm publication, trigger activation, and dependency Use access.

Component Test Matrix

ComponentHappy pathBoundary and failure tests
ConnectorValid credentialExpired, revoked, wrong tenant, missing scope
HTTP actionValid inputs and responseEmpty, malformed, rate-limited, timed out, unauthorized
Script or DSLExpected valuesNull, empty, wrong type, numeric/date boundary
Dynamic resolverOne clear matchNo match, several matches, stale object
SlotClear inferred valueMissing, invalid, ambiguous, recollected
Compound actionAll steps succeedEach step fails, error handling, partial write, return still shaped
Conversation processCorrect collection and resultReordered inputs, cancellation, confirmation rejected
Conversational triggerIntended utteranceNearby capability, vague request, negative examples
WebhookValid signed eventReplay, invalid signature, duplicate, unexpected schema
SchedulePublished, Active, expected run timeInactive trigger, timezone, daylight saving, duplicate or missed run

Verify the Reasoning Boundary

Inspect the values exposed to the conversation:

  • Are keys named for business meaning?
  • Are raw responses, headers, debug values, and unused fields removed?
  • Are compound action internals absent?
  • Does every exposed array need to be in context?
  • Are display instructions short and relevant?
  • Could structured data exceed the direct-response path and require Structured Data Analysis?

If several consecutive action activities expose backend plumbing, move that work into a compound action.

Verify Deterministic Controls

Use deterministic enforcement for:

  • Allowed values and state transitions.
  • Amount and date limits.
  • Authorization and asset access.
  • Required confirmation.
  • Idempotency.
  • Documented error behavior and any custom retry pattern.
  • Record selection from an approved set.
  • Data redaction and field filtering.

Use descriptions and model instructions for language and presentation, not as the only control for consequential rules.

Test Security Boundaries

Verify that secrets exist only in connector or platform credential fields. Remove tokens from example values, URLs, descriptions, logs, screenshots, and copied test payloads.

Confirm the downstream identity can perform only the operations and access only the records required by the plugin.

Require confirmation when a conversational write is destructive, costly, difficult to reverse, or affects another person.

Prevent duplicate writes when a request times out, a webhook retries, or a user submits twice.

Validate provider signatures or challenges before processing event data. Plan for duplicate and out-of-order delivery.

Treat API fields, webhook payloads, uploaded files, and retrieved text as untrusted data. Do not let content override system or business rules.

Use response schemas and output mappings to remove secrets, tokens, private metadata, and unnecessary personal information before values reach conversational context.

Test Retrieval Quality

For conversational plugins, build a small evaluation set:

Phrases that should select the plugin:

  • Direct commands.
  • Questions.
  • Short phrases.
  • Synonyms.
  • Requests with values already supplied.

Update title, description, and examples when classification is wrong. Do not hide an unclear capability behind long prompt instructions.

Observe Production Behavior

After limited launch:

  1. Review conversation and action logs.
  2. Review webhook event and process logs for ambient plugins.
  3. Monitor downstream error rates, rate limits, and token expiration.
  4. Check resolver no-match and disambiguation patterns.
  5. Check which output fields the assistant uses or ignores.
  6. Confirm that provider, webhook, or custom retry behavior does not duplicate writes.
  7. Expand launch access gradually.

Build Along: Test the PurpleSuite Plugin

Run the complete feature-request build as one traceable test:

PurpleSuite end-to-end call trace

Follow one live feature request through the same three API calls and Agent Studio contracts in every chapter.

Shared connector
https://marketplace.moveworks.com
Every API action
Connector adds Bearer PAT. Action adds X-Instance-ID.
Trace one request through every boundary
Record the selected ID and verify that the call sequence is GET list, PATCH update, then GET by ID.
User requestConversationConversational plugin
Move <a live feature request name> to Planned.
Receives
Natural-language intent
Returns
Selected plugin and conversation process
How it fits
The title, description, and triggering examples retrieve the capability. They do not contain an API record ID.
Successful run: actual API call ledger
1GET/api/purple-suite/community/feature_requestsDynamic resolver retrieves live candidates
2PATCH/api/purple-suite/community/feature_requests/{id}Compound action updates the selected record
3GET/api/purple-suite/community/feature_requests/{id}Compound action verifies the stored result
Use a record returned by your own PurpleSuite instance. Save its original status before the write and restore it after mutation testing when appropriate.
  1. Start with a phrase that identifies a live PurpleSuite request and asks to move it to Planned.
  2. Confirm the resolver ran GET /api/purple-suite/community/feature_requests and selected the expected record ID.
  3. Reject confirmation. Verify that no PATCH or verification GET ran and that the record did not change.
  4. Repeat the request and accept confirmation.
  5. Confirm the compound action ran one PATCH /api/purple-suite/community/feature_requests/{id}.
  6. Confirm the next call was GET /api/purple-suite/community/feature_requests/{id} with the same ID.
  7. Verify the stored currentStatus is Planned and the assistant returned only the request name and new status.
  8. Inspect action and process logs for the same run. Verify that no PAT, instance ID, raw headers, or unnecessary response fields appear.
  9. Expire or replace the development credential and confirm the failure is diagnosable without exposing the secret.

Record the utterance, selected ID, expected status, actual status, action instance ID, and result. Your checkpoint is one passing end-to-end trace plus documented failure tests for authentication, no matches, ambiguous matches, invalid status, rejected confirmation, and provider errors.

Definition of Done

  • Every component passes isolated tests.
  • The complete flow passes realistic end-to-end tests.
  • Retrieval selects the plugin and rejects nearby capabilities.
  • System triggers verify and map payloads correctly.
  • Scheduled triggers are published and Active.
  • No system-triggered path assumes meta_info.user.
  • Deterministic rules have deterministic enforcement.
  • Output mappings expose only useful, named values.
  • Conversational confirmation or an explicit ambient approval or handoff protects writes that require human review.
  • Access works through the full dependency chain.
  • Logs support diagnosis without exposing secrets.
  • An owner and credential-rotation process exist.

Continue Learning

Return to the Plugin Building Masterclass overview.