diff options
| author | Kuba Piecuch <jpiecuch@google.com> | 2026-10-03 11:53:17 +0000 |
|---|---|---|
| committer | Tejun Heo <tj@kernel.org> | 2026-10-03 13:23:31 -1000 |
| commit | e6ac89b8b1c12e9104df45a14a26e2dfa96ded06 (patch) | |
| tree | 31e9f8b717945df4cd2452e00f1c865d4a87ed9e /include/linux/sched | |
| parent | cc07fae1de5dbf514a9a3978050ee3cf5fe20d3c (diff) | |
sched_ext: Generate qseq from a per-task counter
finish_dispatch() uses the qseq embedded in p->scx.ops_state to tell
whether the QUEUED instance of a task it's about to claim is the one
scx_bpf_dsq_insert() saw. qseq is generated from rq->scx.ops_qseq, but
the counters of different rqs are independent, so if a task is dequeued
and re-enqueued on a different rq between scx_bpf_dsq_insert() and
finish_dispatch(), the new QUEUED instance can end up with the same
qseq as the old one:
CPU X CPU Z
----- -----
enqueue p on rq A, qseq = N
ops.dispatch()
scx_bpf_dsq_insert(p)
records qseq N
sched_setaffinity(p)
dequeue p from rq A
enqueue p on rq B, qseq = N
finish_dispatch(p, N)
qseq matches, p is claimed
The claim itself is still atomic so the core stays consistent, but an
insert issued for a previous QUEUED instance gets applied to a new one
which the BPF scheduler has just received through ops.enqueue(). This
breaks the guarantee that dispatches targeting a stale instance are
ignored.
Generate qseq from a per-task counter, p->scx.ops_qseq, instead so that
consecutive QUEUED instances of a task never share a qseq regardless of
which rq they're on. The counter is only updated in
scx_do_enqueue_task() with the task's rq locked, so no additional
synchronization is needed, and it fits in an existing hole in struct
sched_ext_entity on 64bit. Remove the now unused rq->scx.ops_qseq.
Never generate qseq 0. NONE and DISPATCHING don't carry a qseq, so
scx_bpf_dsq_insert() on a task in either state records 0. With a
per-task counter, every task's first QUEUED instance would otherwise get
qseq 0 and could be claimed by such an insert. Wrap the counter where
the QSEQ field wraps so that it can't reach a value that shifts to 0 on
32bit either.
Fixes: f0e1a0643a59 ("sched_ext: Implement BPF extensible scheduler class")
Cc: stable@vger.kernel.org # v6.12+
Assisted-by: Claude:claude-opus-5.5
Signed-off-by: Kuba Piecuch <jpiecuch@google.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
Diffstat (limited to 'include/linux/sched')
| -rw-r--r-- | include/linux/sched/ext.h | 1 |
1 files changed, 1 insertions, 0 deletions
diff --git a/include/linux/sched/ext.h b/include/linux/sched/ext.h index 582d7cd4a983..36c04797436c 100644 --- a/include/linux/sched/ext.h +++ b/include/linux/sched/ext.h @@ -207,6 +207,7 @@ struct sched_ext_entity { s32 holding_cpu; s32 selected_cpu; s32 runnable_cpu; /* cpu @p is runnable on, -1 if not */ + u32 ops_qseq; /* protected by rq lock */ struct task_struct *kf_tasks[2]; /* see SCX_CALL_OP_TASK() */ struct list_head runnable_node; /* rq->scx.runnable_list */ |
