I have a Logic App triggered by Service Bus. It calls a backend Service Fabric service through an HTTP action, but the operation takes about 15 minutes. The HTTP action eventually times out and is marked as failed even though the Service Fabric job continues running. What is the best way to handle this timeout and prevent the workflow from reporting a failure incorrectly?
2 Answers
Rather than trying to keep the HTTP request open for 15 minutes, make the operation asynchronous. Have the HTTP endpoint start the Service Fabric job and immediately return a job ID with an accepted response, such as HTTP 202. The Logic App can then poll a status endpoint or wait for a completion event before continuing. This also avoids confusing failures and makes retries safer, especially if starting the same job twice could cause problems.
A shared status store or completion event is generally the right direction. Waiting on a synchronous request for that long can cause problems later if a job hangs or consumes resources while the caller remains blocked.
Depending on the Logic App type and connector configuration, HTTP actions have configurable timeout settings. You can review the action’s timeout configuration and increase it if the platform limit and your reliability requirements allow it. However, for a process that consistently takes around 15 minutes, an asynchronous start-and-monitor design is usually more robust than simply extending the timeout.

Where would you recommend storing the job status? I was considering Blob Storage or a table, but I’m not sure what works best when Service Fabric is running across multiple instances, and I may need approval before adding Blob Storage or SQL.