summaryrefslogtreecommitdiff
path: root/include
diff options
context:
space:
mode:
authorTejun Heo <tj@kernel.org>2026-07-17 22:12:20 -1000
committerTejun Heo <tj@kernel.org>2026-07-19 21:10:57 -1000
commita6ec0b62c589c5b5518e0e3d9eaac6398a82b6f4 (patch)
tree558a843fb34e2aa0a65089518d3d723bd532e804 /include
parent46932bc5fd7ea1156a6742ad1b9306383e0cfb6f (diff)
sched_ext: Hand over cgroups at sub-scheduler enable/disable
Sub-schedulers don't get cgroups yet: every task_group is inited on the root sched and the routing added by the previous patches always resolves to it. Add the handover: an enabling sub-scheduler takes over the cgroups in its subtree and a disabling one returns them to its parent. scx_cgroup_claim_subtree() runs while the sub enables, after the subtree's cgrp->scx_sched's are set and before any task is claimed. It inits each subtree task_group on the sub, exits it from the parent and updates tg->scx.sched. A failed ops.cgroup_init() unwinds the sub-side inits and aborts the enable with the parent untouched. Disabling reverses it with scx_cgroup_return_subtree(): exit each cgroup from the sub, then re-init it on the parent with the current tg->scx.* values, resyncing weight and bandwidth changes made while the sub had it. When a re-init fails, the parent is failed and the remaining task_groups still transfer uninited and get no cgroup ops - the same punting done for tasks. The dying parent's own disable moves them onward. The handover walks include dying but not yet offlined task_groups, the same as root's bulk walks: a removed cgroup keeps hosting scheduling events until its dying tasks finish their final context switches, and its ops.cgroup_exit() must follow the last of them. tg on/offlining is excluded through cgroup_lock(), so either ordering against an rmdir of a subtree cgroup delivers balanced init/exit pairs. Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
Diffstat (limited to 'include')
-rw-r--r--include/linux/sched/ext.h13
1 files changed, 13 insertions, 0 deletions
diff --git a/include/linux/sched/ext.h b/include/linux/sched/ext.h
index a6db5d300f30..78b2f289cb98 100644
--- a/include/linux/sched/ext.h
+++ b/include/linux/sched/ext.h
@@ -298,6 +298,19 @@ static inline bool scx_rcu_cpu_stall(const struct cpumask *stalled_mask) { retur
struct scx_task_group {
#ifdef CONFIG_EXT_GROUP_SCHED
+ /*
+ * The sched this tg is on, NULL if none. SCX_TG_INITED tracks whether
+ * ops.cgroup_init() succeeded on it. When a child sched exits and its
+ * tgs move to the parent, a failed init leaves the tg on the parent
+ * with INITED clear (see scx_cgroup_return_subtree()).
+ *
+ * This is tracked separately from cgrp->scx_sched because the tg
+ * hierarchy can diverge from the cgroup2 hierarchy in both lifetime and
+ * shape. A tg stays online past its cgroup's removal while the
+ * cgrp->scx_sched rewrites visit only live cgroups, leaving a removed
+ * cgroup's pointer stale. The cpu controller can also be mounted on
+ * cgroup1.
+ */
struct scx_sched *sched;
u32 flags; /* SCX_TG_* */