r/AZURE • • Jul 28 '26

Discussion [Discussion+Rant] AML job stuck for 70 days, support denies refund

TL;DR: Azure ML job bypassed its 21-day hard limit and ran for 70 days. A sub-wide Defender for Storage enablement triggered massive scanning costs on the stuck job's continuously appending logs (over than a thousand € for Defender + few hundreds € for the VM). Opened an Enterprise support ticket well within the retention window, but internal MS routing delays caused backend logs to expire. Support now uses the lack of logs to deny any refund, ignoring immutable billing evidence and their own platform limit failure.

I'm dealing with a billing dispute regarding an Azure Machine Learning job and the linked VM + Defender for Cloud costs, and I'm looking for other perspectives and possibly, even some advice with the case and/or dealing with MS Support.

Context:

  • Company: Enterprise environment with Enterprise Agreement and dedicated support.
  • Project: Resource Group dedicated to an MVP started in 2025, average monthly spend is a few hundreds €**.**
  • ML workload: typical job duration is below 30 mins, in rare cases it reached 8-12 hours (max).

Anomaly timeline:

  • Day 1: an Azure ML job starts and seemingly get stuck in a loop. All logs related to this job disappeared from the Storage Account, except for a computeRecord.txt (full path: https://<storageAccountName>.blob.core.windows.net/azureml/ComputeRecord/dcid.<jobName>/) containing two useful info: VMSize (checks out with billing data) and CoreSeconds, confirming the 70 days job duration. NOTE: this file was found (by me) more than a month after the anomaly had ended: until that moment no one (including the MS Support team) had noticed that there was a Job's VM stuck for 70 days behind the Defender for Cloud costs spike.
  • Day 53: colleagues enable Defender for Storage sub-wide with the default 10TB / month cap per Storage Account (a bit too high for a default per Resource cap..?). It continuously scans the AML log files generated by the stuck job, already bloated after almost 2 months of logs, with 3 generating most of the volume (URL with path: https://<storageAccountName>.blob.core.windows.net/azureml/ExperimentRun/dcid.<jobName>/system_logs/... ; files: hosttools-capability.log, lifecycler.log, and metrics-capability.log - biggest one reached 140+ MB, 3 writes per minute -> 3 complete scans per minute).
    • Disclaimer: i acknowledge this is partly on us, as the defaults were accepted and and no paths were excluded.
  • Day 62: 10 TB monthly cap is reached, the Event Grid System Topic created by Defender and linked to the Storage Account shows that the events keep firing, but now they’re being ignored.
  • Day 69: I find out about the issue and having to act quickly to prevent new costs starting the next month (it was a friday afternoon and the next month started with the next monday), I lower the monthly cap to 10 GB and enable Storage Account-level logging/diagnostic logs. Immediately after, the Event Grid System Topic shows the events stop firing: the job abruptly stopped (reason still unknown), which is confirmed by the VM billing. Extra note: the fact that it stopped suddenly when I did some governance operations on Defender and Storage Account makes me think about a platform malfunctioning (just a thought/suspect, obviously I wouldn't use it to push for the refund).
  • Day 75: I open a severity B Support Request under our Enterprise support plan and granting permission for "Advanced diagnostic information".

Incurred costs (using ranges to "anonymize", just in case):

  • VM Compute: 200-500 €
  • Defender for Storage: 1.000€ - 1.500€

The support issue:

  • Official Azure ML docs state a 21-day hard limit for job execution (link): the platform failed to trigger its own timeout. This is the documentation excerpt linked in the original post and in the various emails to MS Support - it's not super clear wheter it's a "hard stop", it just reports it as a "limit":
Resource or Action Maximum limit
Job lifetime 21 days\**1

1 Maximum lifetime is the duration between when a job starts and when it finishes. Completed jobs persist indefinitely. Data for jobs not completed within the maximum lifetime isn't accessible.

  • I initially opened the SR with the Storage team, but due to delayed responses and poor investigation, the specialized AML team was engaged over a month later - even though the AML involvement was clear after at most 2 weeks. By that time, the 30-day internal backend diagnostic logs had expired (see point below).
  • Support is now stonewalling: they refuse to evaluate a refund, claiming that without backend logs, they cannot confirm a platform anomaly, completely ignoring the immutable billing, the computeRecord.txt evidence and the platform's failure to enforce its own 21-day limit.

Questions:

  1. What’s your take on this? Is it reasonable from your POV and experience to push for a full refund (VM + Defender costs), since the massive logs were a direct byproduct of the AML anomaly, or should I solely focus on the VM compute costs (at least the excess over the 21-day limit)? I’d expect at the very least the latter, understanding that the Defender operated correctly and based on "poor" configuration on our side.
  2. How do you successfully escalate past a support tier that behaves like this? Do you have “success stories” or advice regarding similar cases?

Thanks in advance to anyone who finds the willpower to read through this wall of text! 🙂

(I might add an edit or a comment later with a dedicated rant about the abysmal support experience itself - useless pings just to keep the ticket within SLA, completely ignored feedback, repeatedly asking for data I had already provided, and conflicting directions from different teams 🥲*).*

15 Upvotes

19 comments sorted by

6

u/dabrimman Jul 28 '26

Do you have an account manager you can escalate and discuss this with? Nuance isn’t the standard supports strong point.

In my opinion, what seems reasonable is that you pay for 21 days worth of the 69 days worth of usage (30%~). Assuming that it’s documented 21 days is a hard stop for AML.

1

u/Randomusernameeeeee Jul 28 '26 edited Jul 28 '26

There is an Account Manager who has been in CC for the whole time, but (understandably) I doubt he has read anything. I already formally asked for an escalation to the support guy as an answer to the last thread.

This is the documentation excerpt linked in the original post and in the various emails to MS Support:

Resource or Action Maximum limit
Job lifetime 21 days1

1 Maximum lifetime is the duration between when a job starts and when it finishes. Completed jobs persist indefinitely. Data for jobs not completed within the maximum lifetime isn't accessible.

It's not super clear wheter it's a "hard stop", it just reports it as a "limit". The AML specialized support at a certain point has also written this:

[...] the observed activity is tied to core AML runtime components that operate on compute instances or cluster nodes.

These components are part of the platform’s normal operation and are responsible for system monitoring, lifecycle management, and log/telemetry collection. Importantly, they continue to run even when no user workloads are active*, which explains why activity is observed while the compute appears* idle*.*

Which explains why the logs were there and being continuously updated (also in case the job wasn't even running anymore), but not why the VM instance associated to the job (and maybe the job itself) were stuck in execution for 70 days. When I asked clarifications about this, they replied with the usual "logs aren't available, no root cause analysis can be done".

5

u/dizhef Jul 28 '26

That's a toughy, dude. First up, front-line support follows rigid rules - no logs, no ticket. This is an EA/CSAM discussion.

I'd absolutely go in for the ML job costs, but wouldn't put too much time toward it. Their argument will be Defender acted as configured, but you already know the blast radius was worse because of the ML job.

I'm not sure what your overall spend is, but that amount and situation are hard to prosecute without wasting your time.

For my money, I'd ask the CSAM to look at it and weigh in but wouldn't push hard - my time, and others', would outweigh the credit pretty quickly. Unless the MVP is part of something much bigger, in that case I'd be putting pressure on to address the platform failure.

3

u/Randomusernameeeeee Jul 28 '26

Thanks a lot for the reality check. You're right, the recent stonewalling makes complete sense given this rigid "no logs, no ticket" rule. I just wish they had been upfront about it instead of stringing me along and ignoring the billing evidence.

The funny thing is, a few months ago I had a similar issue in the same AML Workspace (we were billed for 20 days for a deallocated VM): support quickly admitted an internal bug and gave us a refund.

Because of that easy win, I went into this SR extremely confident but before I knew it, I had sunk way too much time into it - writing 15+ detailed emails, pulling telemetry and billing data, etc. The fact that I'm 95% sure that this is in fact a platform failure particularly triggered me and I ended up treating it like a personal crusade.

My company lacks FinOps culture and solid governance processes, at least in the subscriptions I work with, and in each sub there are various projects like this one with similar or even greater spend (up to 5-10x): therefore the spike didn't cause much noise, a couple of colleagues barely noticed it at the time but now everyone has already moved on.

I'll see what the Support guy will do after my request and I'll consider bringing this to the CSAM, although I am not sure it is even worth the effort anymore given the context. I would have vastly preferred a strict technical closure on the ticket (even a well-reasoned technical rejection) rather than this bureaucratic dead end, but I'll have to accept this 🥲

Thanks again for the advice!

1

u/matiascoca Jul 31 '26

Had almost this exact shape on Azure Batch in early 2025. Job hit the retention window but a compute pool kept a phantom slot after a control-plane upgrade because the orchestrator lost workload state. Support did exactly what you describe, asked for logs beyond their own 30-day retention then declined the refund when we could not produce them. Ate about 2400 dollars.

What moved the needle was pulling the platform limit doc, the immutable billing line items, and the resource activity log for the parent resource group into a single PDF and escalating through the Enterprise TAM instead of the support ticket queue. TAMs have billing-dispute levers first-line does not. Took six weeks and got a partial refund covering compute past the 21-day mark.

Your case is stronger because the 21-day limit is documented and their failure to enforce it is on them. Push hard on the VM excess past day 21 and go through TAM, not support. Defender you may want to concede early to look reasonable on the shared config point.

1

u/Michal_F Jul 28 '26

I don't see why there should refund you ?

  • Looks like you don't have any subscription budget alerts configured. For any public cloud, daily cost monitoring is a must have. As standard developers, SRE deploy something and really don't check costs daily or monthly.
  • You enabled some feature like Defender for storage without investigation how much data is there written. We had similar case, and cost where much higher, because it was enabled on storage account used for veeam backup target :)
  • For the azure ML, I don't work with this type of cloud resources, but question is if this is your custom job or default behavior. If you configured some ML job, that was running 70 days and nobody noticed something is not right with your process, monitoring, team ...
Just the amount of data written should trigger some anomaly detection, alert.

11

u/SecretaryLife5031 Jul 28 '26

Eh?

There's a hard stop at 21 days. There own platform did not behave as documentation and strong suggests, causing excess billing due to an error on their platform. Yes he should have had budget alerts, and on normal circumstances i would agree, but the nuance is the bug.

3

u/Randomusernameeeeee Jul 28 '26

As i said above, I understand that there were multiple ways to avoid or at least "limit" the damages on our side (path exclusion, budget alerts as you said) - sadly in my company cloud governance and cost control aren't really taken seriously or approached in a standardized and unified way. As a recently hired Cloud Architect with some experience I'm trying to bring these themes up but that's for another rant 🙂

I don't even work with the AML resources except for "high level" governance, two other colleagues managed the jobs execution, but there were plenty of runs before and after this one and each one of them terminated in avg within 30 mins, 8-12 hours at most. Given these numbers, 70 days is clearly an anomaly and I'm pretty sure that it originated by some kind of internal loop (platform side or job's code side).

Sadly, support is right when saying that without these "internal backend logs" not much can be clarified about why this loop started (or finished), but even the fact that it stopped suddenly when I did some governance operations on Defender and Storage Account makes me think about a platform malfunctioning (just a thought/suspect, obviously i wouldn't use it to push for the refund).

I don't know the details about the job's execution and what it does, also because the logs disappeared: I think is somehow expected to happen when the job fails, [other mini rant alert] which feels VERY counterintuitive, why would they remove logs from failed jobs that 99% of the times are those more interesting? But it looks that nothing else but the VM logs were produced by this prolonged execution, maybe that's why nobody noticed anything (also keep in mind this happened within an MVP, thus even more "relaxed" governance and monitoring with respect to the already "relaxed" norm).

2

u/Michal_F Jul 28 '26

In your case I really cannot help as I don't understand Azure ML and other AI tools in azure. I use mostly classic infrastructure resources, and this is easy to monitor and manage. Cloud costs are complicated topic and Cloud Architect will probably not help too much as this more cost optimization/finops related task.

But I recommend to setup budgets + alerts on all subscriptions. We have som estimation how should be the monthly cost and if they are higher than 110% budgetowner and cloud team gets email notification.

I am currently implementing MS finops toolkit just to identify and optimize cost in our environment, but it's probably needed only if you have lots of subscriptions and resources. I like to do statics and reports, so this is ideal for me :)

But we have also cloud cost in our grafana dashboard ... i think after this maybe you can push some ideas to improve your environment.

2

u/Randomusernameeeeee Jul 29 '26

As I said before, governance and cost control really are a mess in the environment I work in: defining a budget for a Subscription wouldn't make sense right now for various reasons, but doing it at a RG scope (at least in some cases) would definitely be possible and useful.

I've been eyeing the finops toolkit for a while, we do have lots of Subs and resources (like 10-15 subs, 1K+ resources) so it looks like it could be helpful. I haven't tried it yet (due to lack of time, costs per se and common costs mgmt not defined), but I'd like to do that in the near future - also I've read it should cost "only" 100-120$ (+ some volume of data-based fees i guess but I'd start from a very restricted scope).

Do you have some advice to share from your experience? For example:

- How are you deploying it? Particularly curious if via IaC since I'm recently leaning towards Bicep (better late than never 😄) and I found out there's an AVM "pattern module".

- How easy is it to customize it, both from a "functional" perspective ("modularity" of some kind?) and for what concerns the scope (e.g. start from 1 sub and increase gradually)?

- What are its biggest benefits? Does it lead you to some quick wins with insights and stuff, or is it "just" a data aggregator with some nice dashboards?

Thanks in advance!

2

u/Michal_F Jul 29 '26 edited Jul 29 '26

Hi, No Problem.

In our case we have new sub for each new workload so this way is much mor easy to setup budget :)

MS finops hub:

  • first note is cost, there are three tiers how this can be deployed. Storage only (minimum cost, like <5 USD/month), with ADX (about 200-250 USD/month but can be reduced to 120 with start/stop automation), with Fabric. The estimate on official web is old and not up to date
  • second note is FOCUS format, you need to configure this cost export to open standard FOCUS format. In this case, this can be limited to subscription or billing account >> all sub. Details depends if you have EA/MCA/Pay as you go ... We have MCA -> manual created export, and planned copy task pipeline. FOCUS format is becoming standard for all big cloud providers. >> https://focus.finops.org/
  • Officially there are no documentation how to deploy it with IaC, but is simple ... code is open source on github, but use stable release tag for deploy >> https://github.com/microsoft/finops-toolkit/tree/dev/src/templates/finops-hub you just create bicep parameter file with required parameters and deploy main.bicep
  • I started with scope all subscriptions, as there is option also to export 13 month of old data in Focus format files.
  • data presentation, they provide PowerBI reports for storage only hub and KQL PowerBI reports for ADX hub. ADX dashboard tempalte ... I learned, it not to hard to modify the reports if you need.
  • For testing you can deploy storage only hub there will be 2x storage accounts and data factory to do some data transformation logic ... One storage account for data export, second is used for data transformation and storage.
  • Then you can connect to the storage account with storage PowerBi report.
  • For me it's data aggregator you can configure to have more than MS default 12-13 months of data. So good to view historic data. We have also AWS data there, but it's not ideal because of currency miss match.
  • Nice view is cost changes per subscription.
  • If you have some special tags like CostCenter, Environment, ... you can also create and make nice dashboards related to this tags.

  • We are also planning to Test AWS Focus dashboard, as just basic dashboards.

  • There are some nice features like recommendation, in this case it can detect some leftover resources. Not connected NSG, disk drives not assigned to any vm, ... and you could create your own queries .. Plus all azure advisor recommendation in one place.

  • ADX uses KQL is the same thing you need to use for Log Analytics workspace, you need to learn this to troubleshoot issues Azure anyway. :)

  • some nice and planned features are AI related, copilot integration ... But this require ADX/Fabric.

It's not perfect tool but evolving, and from MS a good decision to go with open source for this. So you can help, comment and watch what is in progress. But for cost optimization, real issue is how to get right people to the required steps. Many times not many people know what and why something was deployed in some way.

Edit: some typos
Edit2: this was a project I created for me to learn new things, but company hired new FinOps specialist starting in next months, my plan is to move it to prod as their solution. We will see ..

2

u/Randomusernameeeeee Aug 03 '26

Thanks a lot for the super detailed answer, very very interesting! I'll come back to this post for sure when I'll finally pick up the finops toolkit.

I'll have to investigate about which features are included or not in the minimal tier (only Storage) but 99% I'll start with that one. Also, despite having took the exam for the basic FinOps certification some years ago, I don't remember much about Focus but I think that it'd be very useful to review those concepts as well.

And I totally agree with the last part, that's the hardest issue with this job - I hope that at least this toolkit will help me to "frame" the inefficiencies in a way that they are too clear to ignore 😅 and that someone will somehow help me to help the responsible people to take responsibility and action about these, but it's easier said than done.

Good luck with your joint effort with the new FinOps specialist, I'm sure he'll recognize the value of the work you've done 👍

4

u/az987654 Jul 28 '26

"sadly in my company cloud governance and cost control aren't really taken seriously or approached in a standardized and unified way"

then this is the price you pay for not taking it seriously.

2

u/Randomusernameeeeee Jul 28 '26

I understand what you're saying and please believe me when I say that this aspect is maybe the one that frustrates me the most with this job, being a Cloud Architect with some experience with FinOps and having started working for this company one year ago with much higher expectations regarding the "cloud adoption maturity".

That being said, with all the info I provided I think it's clear that something strange happened (most likely a platform malfunctioning with official docs to support this) and together with an overlooked default config and cost mgmt, it caused two different kind of unforecasted expenses of which at least one could be blamed entirely on the CSP.

Imo, it's wrong to put all the blame on me / my company.

1

u/az987654 Jul 28 '26

Why didn't you simply add yourself a calendar reminder on day 21? day 22? no, you waited until day 70+ and now you're wanting to accept any responsibility for the issue.

Why didn't you set a budget alert?

Why didn't you notice your project wasn't finished after you started it?

Why is it everyone else's fault but yours?

2

u/Randomusernameeeeee Jul 28 '26

I'll try to take this as constructive criticism even if it really looks like passive (?) aggressive behaviour.

  1. As stated in other comments, I'm the project's Cloud Architect, I didn't work directly on the AML workspace - 2 other colleagues did. They launched new jobs daily so I'm pretty positive that they'd have noticed a job running for several days - I'm pretty convinced either this "malfunctioning" state hid the job from the UI, or more likely the job actually terminated in a few hours at most but then the job's VM was left stuck in this loop. The VM daily cost was a 3-4€ daily increase in an environment with costs often fluctuating due to other VMs being spun up and deallocated sporadically with no precise patterns. To be clear, the computeRecord.txt log file was found (by me) more than a month after the anomaly had ended: until that moment no one (including the MS Support team) had noticed that there was a Job's VM stuck for 70 days behind the Defender for Cloud costs spike.
  2. I'm incline to answer to this in various ways (maybe all excuses, I agree) - because nobody in my company seems to care much about this, it's not my money, there was no clear budget, I had other stuff to do - but since I'm trying to be(come?) the guy who cares, I acknowledge that part of the blame is on me and I could've and should've done more and better to limit this. However, I still believe that the another (significant) part of the blame is on the platform.
  3. Don't really understand this point - if you're talking about the job, I've already answered above.
  4. See final part of point (2).

3

u/Adures_ Jul 28 '26

It’s platform fault for not enforcing it documented 21 day billing. 

Also if we are being picky, azure alerts are also not be all end all solution, because  1. They only alert about unusual billing happening and do not stop them.  2. They can have massive delay and still not help you avoid to incur massive bill. 

It’s Microsoft fault that their platform does not follow their own documented limits and it’s their fault that shi**y support takes so long that the logs that were available to them expired. 

To OP, my advice is to learn from this mistake.  If Microsoft won’t admit that the additional billings costs   are a result of their problem with platform, you will need probably to pay up. Use it as a reminder and next time you want to deploy some new functionality or service, look to competitors. 

Azure while powerful platform it’s not predictable or helpful if there are any problems, Premium pricing and no support.

 I kid you not I had better experience with cheap Wordpress hosting support than with Azure. 

1

u/Randomusernameeeeee Jul 29 '26

Thanks for your comment, sadly I think "Premium pricing and no support" sums it up perfectly. Where I work I'm completely bound to working with Azure, but I'll keep that in mind for the future - I was already thinking about putting some time into studying AWS.