The vendor feed broke again. Last month it was a timeout. This month the nightly sync stopped halfway through, two records exist at the vendor that shouldn't, and someone is patching rows by hand. Each time the fix was a retry or a try/catch, and each time it held until the next surprise. If your third-party API integration is failing in production and the failures don't repeat the same way twice, the cause is rarely the vendor's bad luck. It is an assumption in your code that the vendor never promised.

Here is a concrete one. The standard resilience handler in Microsoft.Extensions.Http.Resilience retries failed requests up to three times, with a two-second exponential delay and jitter, and by default it does so for every HTTP method. Add it to a client that creates records and you have built a duplicate generator, and you did it with the recommended package. The same Microsoft documentation says so and shows how to switch it off. Most teams don't read that far.

This article names the five assumptions that integrations break, gives each a warning sign you can find in your logs today, and shows the .NET pattern that contains it. It ends with the one job that catches whatever the five patterns miss. It's written for the engineer who owns the integration and the lead who is tired of hearing about it.

Third-Party API Integration Failing? Start With the Five Assumptions

Integrations fail in production for the same reason distributed systems fail anywhere: the other side is a separate machine with its own deployments, its own capacity and its own bad afternoons. Michael Nygard's Release It! (2nd ed., 2018) opens its catalogue of stability antipatterns with integration points for that reason. Every call out of your process is a place where something outside your control can go wrong. What makes it hard is that the failures don't announce themselves. Some throw an exception. The costly ones leave your system and the vendor's disagreeing about what happened, with nothing in the logs.

Here is the model we use to triage a failing integration. Each row is an assumption that holds in a demo and breaks in production.

The assumption What production does Warning sign The fix
A call either worked or it didn't A timeout leaves the outcome unknown Duplicates at the vendor after a slow day No blind retries on writes; look up by your own reference, or use an idempotency key
Capacity is ours to use The vendor throttles, per key, per second or per day 429s that cluster at the start of a batch Pace client-side, honour Retry-After, park long waits on a queue
Success is all or nothing A 200 carries per-item failures; a multi-step flow stops halfway Counts that differ by a few; rows stuck "pending" Per-item results, one state per step, resume from the last completed step
The schema is frozen Fields are renamed, retyped or added, and new status values appear Parse errors after a week with no deploy of yours Tolerant reader, required on what you depend on, contract tests on real payloads
Events arrive once, in order, always Webhooks are duplicated, reordered and sometimes never delivered Vendor and local state disagree and no error was logged Durable inbox, acknowledge fast, fetch current state, reconcile

Read the "warning sign" column as a sorting rule. Failures that throw get fixed, because someone sees them. Failures that leave two systems disagreeing get discovered, by a customer or an auditor, weeks later. Rows one, three and five, and the quiet half of row four, belong mostly to the second group, and they are where the money is lost. The five sections below take the rows in order.

1. Timeouts: The Call Whose Outcome You Don't Know

A timeout isn't a failure. It is the absence of an answer. The vendor may never have seen the request, may have been halfway through it, or may have completed it and lost the response on the way back. The Azure Architecture Center's Retry pattern describes exactly this case: a service processes the request successfully, fails to send the response, and the retry logic sends the request again. If the operation isn't idempotent, it now runs twice.

This is the shape we usually find:

// NAIVE: works until the vendor has a slow afternoon.
public async Task<string> PlaceOrderAsync(Order order)
{
    for (var attempt = 1; attempt <= 3; attempt++)
    {
        try
        {
            using var http = new HttpClient();            // new client per call; 100 s default timeout
            var resp = await http.PostAsJsonAsync(Url, order);  // no key, no CancellationToken
            resp.EnsureSuccessStatusCode();                     // a 400 and a 429 both become "retry"
            return (await resp.Content.ReadFromJsonAsync<OrderDto>())!.Id;
        }
        catch (Exception)                                   // the reason is thrown away
        {
            await Task.Delay(1000);                         // fixed delay, no jitter, ignores Retry-After
        }
    }
    throw new InvalidOperationException("Order failed");
}

Every line there does damage. HttpClient's default timeout is 100 seconds, so a hung vendor holds your request open for well over a minute. The loop retries a POST after a timeout, which is the duplicate-order case. It retries a 400, which will never succeed. And when all three attempts fail, the caller learns "Order failed" and nothing about whether an order exists.

The fix has two parts. Configure the pipeline so it never blind-retries a write, and make the write path handle the unknown outcome explicitly:

builder.Services.AddHttpClient<VendorClient>(c => c.BaseAddress = new("https://api.vendor.example/"))
    .AddStandardResilienceHandler(o =>
        o.Retry.DisableForUnsafeHttpMethods());   // POST, PUT, PATCH, DELETE are not retried by the pipeline

/// <summary>Places an order. Never returns "failed" when the truth is "unknown":
/// after an ambiguous failure it asks the vendor before anyone resends.</summary>
public async Task<PlaceOrderResult> PlaceOrderAsync(Order order, CancellationToken ct)
{
    try
    {
        using var resp = await _http.PostAsJsonAsync("orders", VendorOrderRequest.From(order), ct);

        if (resp.IsSuccessStatusCode)
            return PlaceOrderResult.Confirmed(await ReadOrderAsync(resp, ct));

        if (resp.StatusCode is HttpStatusCode.BadRequest or HttpStatusCode.UnprocessableEntity)
            return PlaceOrderResult.Rejected(await resp.Content.ReadAsStringAsync(ct)); // dead-letter; retrying can't help

        if (resp.StatusCode == HttpStatusCode.TooManyRequests)
            return PlaceOrderResult.Throttled(resp.Headers.RetryAfter?.Delta ?? resp.Headers.RetryAfter?.Date - DateTimeOffset.UtcNow);  // definitely not processed

        return await ResolveUnknownAsync(order, $"HTTP {(int)resp.StatusCode}", ct);          // 5xx: maybe processed
    }
    catch (Exception ex) when (ex is HttpRequestException or TimeoutRejectedException
                               || (ex is TaskCanceledException && !ct.IsCancellationRequested))
    {
        return await ResolveUnknownAsync(order, ex.GetType().Name, ct);
    }
}

private async Task<PlaceOrderResult> ResolveUnknownAsync(Order order, string why, CancellationToken ct)
{
    _log.LogWarning("Order {Reference} outcome unknown ({Why}); checking vendor before any resend",
        order.Reference, why);

    var existing = await FindByReferenceAsync(order.Reference, ct);   // GET: safe to retry, pipeline does
    return existing is null ? PlaceOrderResult.NotPlaced() : PlaceOrderResult.Confirmed(existing);
}

Three things in that code carry the weight. The caller gets a result type with an honest name for each outcome, including Throttled and NotPlaced, so the worker that called it can decide what happens next instead of guessing from an exception. The lookup by order.Reference, a value you generated and sent, is what turns "unknown" into "known". And TimeoutRejectedException is caught explicitly, because Polly's timeout throws that type and not the standard TimeoutException, a detail the Microsoft guide calls out.

One caveat on the lookup: when it finds nothing right after a timeout, that doesn't prove the vendor never received the request, because it may still be processing it. Repeat the lookup after a short delay before anything is resent.

This design needs one of two things from the vendor: a lookup by your own reference, or an idempotency key that makes a repeated write return the original result. If the vendor offers the key, send it and let the pipeline retry writes. Hohpe and Woolf name the receiving side of this in Enterprise Integration Patterns (2003) as the Idempotent Receiver. If the vendor offers neither, you can't safely automate the resend. Route unknown outcomes to a person and tell the vendor's account manager why. Our idempotent message handling guide covers the receiving side of the same problem.

2. Rate Limits: Capacity That Isn't Yours

Vendors throttle. They may do it per key, per second, per day, or per endpoint, and the limit is often lower for the endpoints you call in bulk. Integrations tend to hit it in one of two ways. A backfill or a month-end job sends ten thousand calls in a burst and the daily sync that shares the key starts failing. Or a retry policy meets a vendor that is already struggling, and the retries become the load.

Know what the protocol promises. A 429 response from RFC 6585 may include a Retry-After header, and the RFC doesn't require it. When it is present, RFC 9110 says it holds either a number of seconds or an HTTP date, so parse both. The .NET standard handler already retries 429 and 408 along with server errors, and its ShouldRetryAfterHeader option defaults to true, so a vendor that sends the header gets obeyed.

The catch is the budget. The handler's total timeout defaults to 30 seconds and covers every attempt including the waits between them. A vendor that answers "retry in 60 seconds" can't be waited out inside a 30-second request. Two changes fix it. First, pace your own calls below the vendor's published limit with a client-side limiter (System.Threading.RateLimiting ships a token bucket) so 429s are the exception. The handler's own rate limiter only caps concurrent requests, at 1,000 by default, which protects your process and says nothing about the vendor. Second, when a 429 does arrive with a long wait, don't hold the request. Return Throttled(delay), as the code above does, and let a queue redeliver the work after that delay. A backfill then becomes a slow, steady drain, and your daily sync keeps its share of the limit.

3. Partial Success: The 200 That Hides a Failure

Two designs cause most of these incidents, and both are common vendor behaviour. The first is the batch endpoint that returns 200 for the request and reports failures per item in the body. Code that checks IsSuccessStatusCode and moves on has just marked every failed item as sent. The second is the multi-step flow: create the customer, attach the payment method, confirm the subscription. Your process dies between steps two and three and the vendor holds a half-built object that nothing in your database knows about.

The warning sign is arithmetic. Records sent minus records confirmed is a small, nonzero number that nobody can explain, or rows sit in "pending" for days. The fix is structural. Read every per-item result, and store the outcome per item, not per batch. For multi-step flows, give each step its own persisted state keyed on your reference, so a restart resumes at the step that didn't complete instead of starting over or giving up. And keep the state machine narrow: Requested, Confirmed, Rejected, NeedsReview is usually enough. Anything that has been in Requested longer than it should be is a finding for the reconciliation job covered below.

4. Format Drift: The Contract Nobody Signed

Vendors change payloads: a field is renamed, a number becomes a string, a new status value appears. Often they don't announce it. The integration then fails in one of two ways, and the quiet one is worse. The loud one is a type change: a number that becomes a string makes System.Text.Json throw. The quiet one is a rename. Properties in the payload that your type doesn't declare are ignored by default, which is the right default. But a property your type does declare that the vendor renamed simply arrives as null, 0 or Guid.Empty, and the rest of your code runs on it happily.

The pattern is Martin Fowler's Tolerant Reader (2011): take only what you need, from one place. We add one rule of our own: be strict about the parts you can't do without. In C#:

/// <summary>Only what we depend on. Unknown fields are ignored on purpose.</summary>
public sealed record VendorOrder
{
    public required string Id        { get; init; }   // renamed or removed: fails loudly, not Guid.Empty
    public required string Reference { get; init; }
    public required string Status    { get; init; }   // a string, so a new value can't throw
    public decimal? Total            { get; init; }   // optional: absence is a business decision
}

public static OrderState MapStatus(string status, ILogger log) => status switch
{
    "pending"  => OrderState.Requested,
    "accepted" => OrderState.Confirmed,
    "rejected" => OrderState.Rejected,
    _ => Unrecognized(status, log)
};

private static OrderState Unrecognized(string status, ILogger log)
{
    log.LogWarning("Vendor sent unrecognized order status {Status}; routing to review", status);
    return OrderState.NeedsReview;   // visible, countable, and no data lost
}

required makes a missing member a deserialization error, so the rename shows up as a failure at the boundary on day one. The status mapping turns a new vendor value into a warning and a review queue instead of an exception that takes down the batch. Then add two habits around it. When deserialization does fail, store the raw payload in a quarantine table with a correlation id before you surface the error, because the payload is your only evidence when you write to the vendor. And keep a contract test that replays captured real responses with JsonSerializerOptions.UnmappedMemberHandling set to Disallow, which .NET 8 and later support. That test fails when the vendor adds a field, so new fields appear in your pipeline before they appear in production, while the production type keeps ignoring them.

5. Lost Webhooks: The Event That Never Arrived

Webhooks look like the easy half of an integration: the vendor calls you. In production they are duplicated, reordered, delivered while your endpoint is deploying, and occasionally not delivered at all once the vendor's retry window closes. Retry windows differ by vendor and are documented per vendor, so treat the exact policy as something to read, not assume. The design consequence is the same in every case: your endpoint must be cheap and durable, and it can't be the only path by which you learn what happened.

app.MapPost("/webhooks/vendor", async (HttpRequest req, IWebhookInbox inbox,
    IVendorSignature signature, ILogger<Program> log, CancellationToken ct) =>
{
    using var buffer = new MemoryStream();
    await req.Body.CopyToAsync(buffer, ct);
    var body = buffer.ToArray();

    if (!signature.IsValid(req.Headers, body))       // vendor-specific scheme; constant-time compare
        return Results.Unauthorized();

    var evt = VendorEventEnvelope.Parse(body);        // envelope only: id and type, nothing else
    var added = await inbox.TryAddAsync(evt.Id, evt.Type, body, ct);   // unique index on evt.Id
    if (!added)
        log.LogInformation("Duplicate webhook {EventId} ignored", evt.Id);

    return Results.Accepted();                       // acknowledge only after the durable write
});

The handler does four things and nothing else. It verifies authenticity, writes the raw event to a durable inbox, deduplicates on the vendor's event id with a unique index, and returns a 2xx. All business logic runs afterwards in a worker that reads the inbox. That keeps the response fast enough that the vendor doesn't see a timeout and retry, and it means a bug in your processing code costs you a replay, not a lost event.

Then change what the worker trusts. Treat the payload as a hint, and fetch the current state of the object from the vendor's API before acting on it. That makes duplicates harmless and reordering irrelevant: a stale "order shipped" event that arrives after "order cancelled" can't overwrite the truth, because the worker asked for the truth. The cost is one API call per event, so count it against your rate limit from section two.

Reconciliation: The Backstop for Everything Else

The five patterns contain the failures you predicted. Reconciliation catches the ones you didn't. It is a scheduled job that compares what your system believes it sent, in your own ledger, against what the vendor says it holds, then sorts the differences into three buckets: missing at the vendor, missing locally, and present in both but different. Safe classes heal themselves, such as an item that is missing at the vendor with no unknown outcome, which gets requeued. Everything else opens a ticket with both records attached.

Two measurements make it useful. First, the age of the oldest unreconciled item, which is the single metric that tells you how long a silent failure can live in your system before someone sees it. Second, the daily count of items the job had to repair. A steady low number is healthy. A rising one is an integration degrading, and you'll see it days before a customer does. The same discipline sits underneath auditable backend systems, where completeness checks against an external source are how you prove nothing was dropped.

What This Design Doesn't Solve

  • A vendor with no lookup and no idempotency key. If you can't ask "did this happen?" and can't repeat a write safely, the unknown outcome can't be resolved in code. Automate detection and route the case to a person.
  • A vendor outage. A circuit breaker makes the failure fast and stops you hammering a struggling dependency. It doesn't make the vendor available. Decide in advance what degrades: queue and drain later, or show the user a clear message.
  • Reconciliation cost. The job spends API quota and needs a vendor endpoint that can list by time window. If the vendor offers only per-item lookups at a low limit, reconcile a sample daily and the full set weekly.
  • Ordering guarantees. Fetching current state fixes stale events for single objects. It doesn't help when a business rule depends on the order of events across objects. That needs its own sequencing design.
  • The vendor's own bugs. Sometimes the fix is a support ticket. What the patterns above buy you is the evidence: a quarantined payload, a request correlation id and a timeline.
Common Questions

Frequently Asked Questions

Why does a third-party API work in testing but fail in production?
Test traffic is small, sequential and clean. Production adds volume, concurrency, real-world data and the vendor's own deployments, and each one breaks an assumption the test never exercised: that a call either worked or didn't, that capacity is unlimited, that success is all-or-nothing, that the schema is frozen, and that events arrive once, in order and always. The code that passes in a sandbox usually fails on the first slow response, the first throttled burst or the first field the vendor renames.
Should I retry a POST request that timed out?
Not blindly. A timeout means the outcome is unknown: the vendor may have processed the request and the response was lost. Retry only if the vendor supports an idempotency key, or if you first look the result up by your own reference. The standard .NET resilience handler retries every HTTP method by default, so for a client that creates records call DisableForUnsafeHttpMethods on its retry options and handle the unknown outcome in your own code.
How do I handle 429 Too Many Requests in .NET?
Treat a 429 as 'not processed, come back later'. The Retry-After header is optional for servers to send, and when it is present it holds either a number of seconds or an HTTP date, so honour it when it is there and back off with jitter when it isn't. The standard resilience handler retries 429 responses and uses Retry-After by default, but its 30-second total timeout covers every attempt, so a longer wait belongs on a queue rather than inside the request. Also pace your own calls below the vendor's limit so you are not relying on 429s to find it.
How do you make webhook handling reliable?
Verify the signature, write the raw event to a durable inbox with a unique index on the event id, return a 2xx response, and process the event afterwards in a worker. Treat the payload as a hint and fetch the current state from the vendor's API, because events can be duplicated and can arrive out of order. Then add a scheduled reconciliation job, because a vendor's delivery retries eventually stop and an event that never arrived leaves no error to alert on.
How do you detect that a vendor changed their API without telling you?
Model only the fields you use and mark the essential ones as required, so a renamed or removed field fails loudly instead of binding a default value. Keep status-like fields as strings mapped through a switch, so a new value is logged and routed to review instead of throwing. Replay captured real payloads in a contract test with unmapped members disallowed, which surfaces new fields in your pipeline before production does. In production, alert on the parse-failure and unrecognised-value rate, not only on HTTP errors.

When to Bring in External Help

You can usually fix one of these in a sprint. Several at once, on an integration that is also carrying revenue or filings, is where an outside diagnosis pays for itself: the failure modes interact, and fixing the loud ones first hides the silent ones. We see this on integrations where a partner's side is still changing. In our 1099 reporting automation work, every filing goes out over an HTTPS API to a partner whose systems we don't control, and the first digital-asset cycle went through a deliberate manual fallback while the partner's new path was finished. Correctness outranked automation, and that is the same trade the patterns above make.

If your backend shows any of the patterns above, our Discovery Sprint is a two-week diagnostic that tells you exactly what to fix and in what order. For an integration that means mapping each call to the assumption it makes, finding the unknown outcomes and the silent divergences, and giving you a ranked plan. The API integration rescue work picks up from there. If the failure you are chasing looks like slowness rather than errors, start with our guide to diagnosing a slow .NET API, whose last step covers downstream dependencies.

Free Resource Before a broader conversation, run your system against our free 12-Point Backend Health Checklist. Its data-integrity and observability checks are the ones an integration depends on.
Back to Insights