[Swarming] Reconsiders swarming tasks for remote execution - #5476
Conversation
dylanjew
left a comment
There was a problem hiding this comment.
mostly LGTM. just a couple questions
PauloVLB
left a comment
There was a problem hiding this comment.
Just a heads up on something this unblocks.
When swarming can't take a task (queue full, API error), the gate hands it over to gcp_batch:
clusterfuzz/src/clusterfuzz/_internal/remote_task/remote_task_gate.py
Lines 192 to 195 in d6d9c25
And batch resolves the specs for the whole list before creating anything, so one task whose platform isn't in the batch config kills the entire chunk:
Saw it in dev with ValueError: No mapping for ANDROID_EMULATOR-PREEMPTIBLE-UNPRIVILEGED:
It doesn't hurt much in dev because gcp_batch_jobs_frequency is 0.2, so the slice is a prefix of the list and swarming tasks are appended last, meaning most cycles only send regular tasks to batch. In prod the frequency is 1.0, so the slice is always the whole list: one bad swarming task fails the whole cycle, and since nothing gets acked it comes back a few minutes later and fails again.
Not asking to change this PR, but I think we should make batch skip the tasks it can't map instead of raising for the whole list before we let this run in prod.
This sounds like something we should fix, but are you thinking it's not a blocker because the frequency of swarming not being able to take the task and falling back to gcp_batch is relatively low? If it continously fails and doesn't ack that sounds concerning We can discuss in a separate bug/PR, but I also would expect failed swarming tasks to get retried on swarming, and failed batch tasks to get retried on batch, rather than skipping any retries for swarming tasks that we can't map to batch. |
|
Thanks for your findings paulo, Will leave this PR as it is (approved but not merged) and will stack a new one top of it changes so that unscheduled swarming tasks doesn't overthrows the whole batch of tasks as invalid. Regarding the 28 tasks, i am positive said tasks are because the job no longer exists in the db, because currently the job-exporter is broken and besides breaking issues(like missing spec or config) would show more errors. |
785e5a3 to
eba0823
Compare
eba0823 to
6586f1e
Compare
As found out in the parent PR: #5476 In the scheduler, when a task fails to get scheduled in swarming, it then tries to go on to other backends, so if for any reason(e.g. the platform is not supported in batch per the batch config), then the whole slice of batch tasks get rejected. So for this PR, we make it so swarming tasks are only ever tried on swarming, and other tasks continue to be tried on other backends. Also when making this changes i thought of an edge case not considered previously, what if there was a swarming task in the queue & then we disable the flag? Then the scheduled would mark the task as not swarming & would try it on other backends, and again this could show errors if the platform is not supported there. So i changed the validations. With this changes task now appear in swarming! <img width="704" height="511" alt="image" src="https://github.com/user-attachments/assets/df9665e9-7257-4db6-a38d-22384b6eedcc" /> ## Changes - `swarming/__init__.py`: Now we allow to validate if a given job is swarming or not without considering the flag. - `remote_task_gate.py`: Splits tasks in swarming and other, swarming tasks get redirected to the swarming service and the unscheduled tasks from this service are not added back again with the other tasks ## Tests - A bunch of Tests + refactoring some tests to support this new edge cases. - Left this changes running in dev for 24 hours, and no longer saw the error that canceled the whole slice of batch tasks if the job platform didn't existed for batch, I only saw this error: <img width="1572" height="563" alt="image" src="https://github.com/user-attachments/assets/cbbeb1a1-dad9-42a3-96fb-0da48ae7b7a2" /> At first sight i also thought that a swarming job ended up being considered for batch, but we can ignore that because the real root cause to even being it considered for batch is because the job doesn't exist! So, this kind of errors doesn't appear because of a swarming task was considered for batch, it appears because a non existen job was considered for batch, im positive this error happens because of the problems between the job-exporter and because that's a job that i manually inserted/scheduled to test the parent ticket.


We laid the ground for swarming scheduling quite some time ago, but due to other blockers(i.e. android fuzzing issues), we never really tested the scheduling in prod.
The issue here is that
batch_service.is_remote_task()only returns true for tasks whose platform has a matching config in thebatch.yamlfile, i suspect we didn't saw this before because until recently, android emulator fuzzing, was scheduled as aLINUXplatform, which is considered in batch config, were asANDROID_EMULATORis notChanges
Tests
Swarming job Before:

After: With this small change swarming task are now successfully redirected to the utask_main queue:

So that they are processed by the utask_main scheduler