My Logic App is triggered by Service Bus and calls an HTTP endpoint that starts a Service Fabric job. The backend operation usually takes about 15 minutes, so the HTTP action times out and reports failure even though the job continues running. Is there a way to increase the timeout, or what is the recommended pattern for handling this long-running operation safely?
2 Answers
Rather than simply increasing the timeout, make the operation asynchronous. Have the HTTP endpoint start the Service Fabric job and immediately return a 202 response with a job ID. The Logic App can then poll a status endpoint or wait for a completion event before continuing. This avoids reporting a failure while the job is still running and makes retries safer, especially if starting the same job twice would cause problems.
A shared status store or completion event is the right direction. Waiting synchronously for 15 minutes can cause more problems if the service hangs or consumes resources while the request remains open.
HTTP actions in Logic Apps have configurable timeout limits, so you can review the action’s timeout setting. However, for an operation that consistently takes around 15 minutes, an asynchronous design with a job ID and status tracking is generally more reliable than keeping the HTTP request open.

Where should the job status be stored? I could use a table, blob storage, or SQL, but I’m unsure which option fits when the Service Fabric job runs across multiple instances.