tmp_suning_uos_patched

History

Ingo Molnar 67f9a619e7 [PATCH] sched: fix SMT scheduler latency bug William Weston reported unusually high scheduling latencies on his x86 HT box, on the -RT kernel. I managed to reproduce it on my HT box and the latency tracer shows the incident in action: _------=> CPU# / _-----=> irqs-off \| / _----=> need-resched \|\| / _---=> hardirq/softirq \|\|\| / _--=> preempt-depth \|\|\|\| / \|\|\|\|\| delay cmd pid \|\|\|\|\| time \| caller \ / \|\|\|\|\| \ \| / du-2803 3Dnh2 0us : __trace_start_sched_wakeup (try_to_wake_up) .............................................................. ... we are running on CPU#3, PID 2778 gets woken to CPU#1: ... .............................................................. du-2803 3Dnh2 0us : __trace_start_sched_wakeup <<...>-2778> (73 1) du-2803 3Dnh2 0us : _raw_spin_unlock (try_to_wake_up) ................................................ ... still on CPU#3, we send an IPI to CPU#1: ... ................................................ du-2803 3Dnh1 0us : resched_task (try_to_wake_up) du-2803 3Dnh1 1us : smp_send_reschedule (try_to_wake_up) du-2803 3Dnh1 1us : send_IPI_mask_bitmask (smp_send_reschedule) du-2803 3Dnh1 2us : _raw_spin_unlock_irqrestore (try_to_wake_up) ............................................... ... 1 usec later, the IPI arrives on CPU#1: ... ............................................... <idle>-0 1Dnh. 2us : smp_reschedule_interrupt (c0100c5a 0 0) So far so good, this is the normal wakeup/preemption mechanism. But here comes the scheduler anomaly on CPU#1: <idle>-0 1Dnh. 2us : preempt_schedule_irq (need_resched) <idle>-0 1Dnh. 2us : preempt_schedule_irq (need_resched) <idle>-0 1Dnh. 3us : __schedule (preempt_schedule_irq) <idle>-0 1Dnh. 3us : profile_hit (__schedule) <idle>-0 1Dnh1 3us : sched_clock (__schedule) <idle>-0 1Dnh1 4us : _raw_spin_lock_irq (__schedule) <idle>-0 1Dnh1 4us : _raw_spin_lock_irqsave (__schedule) <idle>-0 1Dnh2 5us : _raw_spin_unlock (__schedule) <idle>-0 1Dnh1 5us : preempt_schedule (__schedule) <idle>-0 1Dnh1 6us : _raw_spin_lock (__schedule) <idle>-0 1Dnh2 6us : find_next_bit (__schedule) <idle>-0 1Dnh2 6us : _raw_spin_lock (__schedule) <idle>-0 1Dnh3 7us : find_next_bit (__schedule) <idle>-0 1Dnh3 7us : find_next_bit (__schedule) <idle>-0 1Dnh3 8us : _raw_spin_unlock (__schedule) <idle>-0 1Dnh2 8us : preempt_schedule (__schedule) <idle>-0 1Dnh2 8us : find_next_bit (__schedule) <idle>-0 1Dnh2 9us : trace_stop_sched_switched (__schedule) <idle>-0 1Dnh2 9us : _raw_spin_lock (trace_stop_sched_switched) <idle>-0 1Dnh3 10us : trace_stop_sched_switched <<...>-2778> (73 8c) <idle>-0 1Dnh3 10us : _raw_spin_unlock (trace_stop_sched_switched) <idle>-0 1Dnh1 10us : _raw_spin_unlock (__schedule) <idle>-0 1Dnh. 11us : local_irq_enable_noresched (preempt_schedule_irq) <idle>-0 1Dnh. 11us < (0) we didnt pick up pid 2778! It only gets scheduled much later: <...>-2778 1Dnh2 412us : __switch_to (__schedule) <...>-2778 1Dnh2 413us : __schedule <<idle>-0> (8c 73) <...>-2778 1Dnh2 413us : _raw_spin_unlock (__schedule) <...>-2778 1Dnh1 413us : trace_stop_sched_switched (__schedule) <...>-2778 1Dnh1 414us : _raw_spin_lock (trace_stop_sched_switched) <...>-2778 1Dnh2 414us : trace_stop_sched_switched <<...>-2778> (73 1) <...>-2778 1Dnh2 414us : _raw_spin_unlock (trace_stop_sched_switched) <...>-2778 1Dnh1 415us : trace_stop_sched_switched (__schedule) the reason for this anomaly is the following code in dependent_sleeper(): /* * If a user task with lower static priority than the * running task on the SMT sibling is trying to schedule, * delay it till there is proportionately less timeslice * left of the sibling task to prevent a lower priority * task from using an unfair proportion of the * physical cpu's resources. -ck / [...] if (((smt_curr->time_slice (100 - sd->per_cpu_gain) / 100) > task_timeslice(p))) ret = 1; Note that in contrast to the comment above, we dont actually do the check based on static priority, we do the check based on timeslices. But timeslices go up and down, and even highprio tasks can randomly have very low timeslices (just before their next refill) and can thus be judged as 'lowprio' by the above piece of code. This condition is clearly buggy. The correct test is to check for static_prio _and_ to check for the preemption priority. Even on different static priority levels, a higher-prio interactive task should not be delayed due to a higher-static-prio CPU hog. There is a symmetric bug in the 'kick SMT sibling' code of this function as well, which can be solved in a similar way. The patch below (against the current scheduler queue in -mm) fixes both bugs. I have build and boot-tested this on x86 SMT, and nice +20 tasks still get properly throttled - so the dependent-sleeper logic is still in action. btw., these bugs pessimised the SMT scheduler because the 'delay wakeup' property was applied too liberally, so this fix is likely a throughput improvement as well. I separated out a smt_slice() function to make the code easier to read. Signed-off-by: Ingo Molnar <mingo@elte.hu> Signed-off-by: Andrew Morton <akpm@osdl.org> Signed-off-by: Linus Torvalds <torvalds@osdl.org>		2005-09-10 10:06:23 -07:00
..
irq	[PATCH] CHECK_IRQ_PER_CPU() to avoid dead code in __do_IRQ()	2005-09-07 16:57:29 -07:00
power	Merge linux-2.6 with linux-acpi-2.6	2005-09-08 01:45:47 -04:00
acct.c	[PATCH] largefile support for accounting	2005-09-07 16:57:31 -07:00
audit.c	[NETLINK]: Add "groups" argument to netlink_kernel_create	2005-08-29 16:01:11 -07:00
auditsc.c
capability.c
compat.c
configs.c
cpu.c
cpuset.c	[PATCH] cpuset semaphore depth check deadlock fix	2005-09-10 10:06:21 -07:00
crash_dump.c
dma.c
exec_domain.c
exit.c	[PATCH] files: files struct with RCU	2005-09-09 13:57:55 -07:00
extable.c
fork.c	[PATCH] files: files struct with RCU	2005-09-09 13:57:55 -07:00
futex.c	[PATCH] futex: remove duplicate code	2005-09-07 16:57:33 -07:00
intermodule.c	[PATCH] introduce and use kzalloc	2005-09-07 16:57:45 -07:00
itimer.c
kallsyms.c
Kconfig.hz
Kconfig.preempt
kexec.c
kfifo.c
kmod.c
kprobes.c	[PATCH] kprobes: fix bug when probed on task and isr functions	2005-09-07 16:58:01 -07:00
ksysfs.c
kthread.c
Makefile	[PATCH] spinlock consolidation	2005-09-10 10:06:21 -07:00
module.c	[PATCH] flush icache early when loading module	2005-09-07 16:57:26 -07:00
panic.c
params.c	[PATCH] introduce and use kzalloc	2005-09-07 16:57:45 -07:00
pid.c
posix-cpu-timers.c
posix-timers.c	[PATCH] fix send_sigqueue() vs thread exit race	2005-09-07 16:57:33 -07:00
printk.c	[PATCH] Provide better printk() support for SMP machines	2005-09-07 16:57:18 -07:00
profile.c
ptrace.c	[PATCH] remove duplicated code from proc and ptrace	2005-09-07 16:57:43 -07:00
rcupdate.c	[PATCH] files: rcuref APIs	2005-09-09 13:57:54 -07:00
resource.c	[PATCH] introduce and use kzalloc	2005-09-07 16:57:45 -07:00
sched.c	[PATCH] sched: fix SMT scheduler latency bug	2005-09-10 10:06:23 -07:00
seccomp.c
signal.c	[PATCH] fix send_sigqueue() vs thread exit race	2005-09-07 16:57:33 -07:00
softirq.c	[PATCH] revert bogus softirq changes	2005-07-30 10:49:59 -07:00
softlockup.c	[PATCH] detect soft lockups	2005-09-07 16:57:17 -07:00
spinlock.c	[PATCH] spinlock consolidation	2005-09-10 10:06:21 -07:00
stop_machine.c
sys_ni.c	[PATCH] remove sys_set_zone_reclaim()	2005-08-01 10:03:56 -07:00
sys.c	[PATCH] remove a redundant variable in sys_prctl()	2005-09-07 16:57:32 -07:00
sysctl.c	[NET]: Fix sparse warnings	2005-08-29 16:01:32 -07:00
time.c	[PATCH] clean up inline static vs static inline	2005-07-27 16:26:20 -07:00
timer.c	[PATCH] optimize writer path in time_interpolator_get_counter()	2005-09-07 16:57:24 -07:00
uid16.c
user.c
wait.c
workqueue.c	[PATCH] introduce and use kzalloc	2005-09-07 16:57:45 -07:00