Skip to main content

Why Automatic Retries Can Create Duplicate Orders

"

Execution and delivery failure

A retry mechanism solves temporary failure only if the action can be repeated safely. If the first request succeeded but the response was lost, a retry can create the same order, payment or record again.

What the symptom actually proves

A retry mechanism solves temporary failure only if the action can be repeated safely. If the first request succeeded but the response was lost, a retry can create the same order, payment or record again.

A useful diagnosis begins by separating what is directly observed from what is only suspected. Execution status, HTTP codes, model output, approval state and destination records are evidence. Statements such as ""the API is broken"" or ""the model ignored the prompt"" remain hypotheses until the workflow trace supports them.

Evidence to collect before changing the workflow

Capture the smallest set of evidence that lets you reconstruct the incident. Redact secrets, personal data and tokens before sharing screenshots or logs.

  • Original request or event ID
  • Destination record created during the first attempt
  • Timeout or connection error details
  • Retry count and exact action repeated

Keep timestamps and stable identifiers wherever possible. They let you correlate the source event, workflow execution and downstream side effect without relying on memory.

Likely failure paths

Do not treat every failure as retryable. Authentication errors, validation errors, duplicate effects, security failures and transient dependency problems require different responses.

  • The first write succeeded but the client never received confirmation
  • The workflow retries an entire sequence instead of the failed safe portion
  • The destination has no unique business key
  • Manual replay occurs after an automatic retry already succeeded

Resolution sequence

  1. Determine whether the first attempt created the business object
  2. Use a stable idempotency or business key on mutating requests
  3. Retry from the smallest safe boundary
  4. Before manual replay, inspect destination state and execution history

If one step requires a broader permission, destructive action, credential exposure or production-data change, move that step into an explicit review or controlled test environment rather than broadening access just to make the run succeed.

Prevention design

  • Design every retryable write for duplicate suppression
  • Store processed event IDs
  • Prefer find-or-create or upsert patterns when supported
  • Document which steps are safe to retry and which require review

The prevention layer should make the next incident easier to detect and cheaper to contain. That normally means stable identifiers, bounded retries, observable execution state, explicit ownership and guardrails around consequential actions.

Verify the fix

A green run is not enough. Verification should repeat the original failure condition and check that no hidden duplicate, unsafe action or stale downstream state remains.

  • Force a controlled timeout after a successful destination write
  • Confirm the retry returns or reuses the existing business result
  • Confirm exactly one order or record exists
  • Review logs to ensure the duplicate was suppressed intentionally

Decision table

QuestionIf yesIf no
Can you reproduce the same failure with a known input?Use that case as the primary regression test.Preserve logs and monitor until the condition recurs or isolate a safe equivalent.
Did a business-side effect already occur?Check idempotency and destination state before replay.Retry may be safer, but only after classifying the error.
Does the fix require more permissions?Reconsider the design and apply least privilege.Keep the current security boundary.
Can monitoring detect recurrence?Deploy with an owned alert path.Add observability before calling the issue closed.

Related reliability guides

Sources and scope

These sources support the platform behavior, reliability controls and security boundaries used in this guide. Platform behavior, limits and interfaces can change, so confirm the current documentation before changing a production workflow.

How this guide was built

This page follows the site methodology: start from a reproducible operational problem, use primary platform or security documentation for technical claims, separate evidence from inference, recommend bounded corrective actions, and end with a verification test. The page intentionally avoids hidden SEO text, invented benchmarks and unsupported guarantees.

"

Comments

Popular posts from this blog

Self-Hosted n8n Keeps Crashing: What Evidence to Collect

" Operations and recovery A self-hosted workflow platform adds infrastructure failure modes to workflow failure modes. Changing container settings, database state and workflow logic at the same time makes diagnosis harder. Resolution rule Change one boundary at a time. Preserve the failing evidence, apply the smallest safe correction, then replay the known case and check the business effect as well as the technical run status. Contents What the symptom actually proves Evidence to collect before changing the workflow Likely failure paths Resolution sequence Prevention design Verify the fix Decision table Sources and scope What the symptom actually proves A self-hosted workflow platform adds infrastructure failure modes to workflow failure modes. Changing container settings, database state and workflow logic at the same time makes diagnosis harder. A useful diagnosis begins by separating what is directly observed from what is only suspected. Execution status, HTTP co...

Contact Feedback

" Contact and feedback Useful technical feedback includes enough evidence to locate a problem while protecting credentials, customer data and confidential workflow information. Contents What to include in technical feedback What not to send How to use available site feedback channels What to include in technical feedback The URL of the page you are commenting on. The specific statement, step or link that appears wrong or outdated. The current platform and version where relevant. A link to current primary documentation if you have one. A sanitized example that does not contain credentials or personal data. What not to send API keys, passwords, access tokens or cookies. Customer names, email addresses, financial data or confidential records. Production database exports. Private webhook URLs or credential-bearing screenshots. How to use available site feedback channels Use the feedback or contact channel made available on this Blogger site when enabled. If article comme...