pid_ns: (BUG 11391) change ->child_reaper when init->group_leader exits

We don't change pid_ns->child_reaper when the main thread of the
subnamespace init exits.  As Robert Rex <> pointed
out this is wrong.

Yes, the re-parenting itself works correctly, but if the reparented task
exits it needs ->parent->nsproxy->pid_ns in do_notify_parent(), and if the
main thread is zombie its ->nsproxy was already cleared by

Introduce the new function, find_new_reaper(), which finds the new
->parent for the re-parenting and changes ->child_reaper if needed.  Kill
the now unneeded exit_child_reaper().

Also move the changing of ->child_reaper from zap_pid_ns_processes() to
find_new_reaper(), this consolidates the games with ->child_reaper and
makes it stable under tasklist_lock.

* the child reaper process (ie "init") in our pid
* space.
static struct task_struct *find_new_reaper(struct task_struct *father)
struct pid_namespace *pid_ns = task_active_pid_ns(father);
struct task_struct *thread;
thread = father;
while_each_thread(father, thread) {
if (thread->flags & PF_EXITING)
if (unlikely(pid_ns->child_reaper == father))
pid_ns->child_reaper = thread;
return thread;
if (unlikely(pid_ns->child_reaper == father)) {
if (unlikely(pid_ns == &init_pid_ns))
panic("Attempted to kill init!");
* We can not clear ->child_reaper or leave it alone.
* There may by stealth EXIT_DEAD tasks on ->children,
* forget_original_parent() must move them somewhere.
pid_ns->child_reaper = init_pid_ns.child_reaper;
return pid_ns->child_reaper;
static void forget_original_parent(struct task_struct *father)
struct task_struct *p, *n, *reaper = father;
struct task_struct *p, *n, *reaper;
reaper = find_new_reaper(father);
* First clean up ptrace if we were using it.
ptrace_exit(father, &ptrace_dead);
do {
reaper = next_thread(reaper);
if (reaper == father) {
reaper = task_child_reaper(father);
} while (reaper->flags & PF_EXITING);
list_for_each_entry_safe(p, n, &father->children, sibling) {
p->real_parent = reaper;
if (p->parent == father) {
static inline void check_stack_usage(void) {}
static inline void exit_child_reaper(struct task_struct *tsk)
if (likely(tsk->group_leader != task_child_reaper(tsk)))
if (tsk->nsproxy->pid_ns == &init_pid_ns)
panic("Attempted to kill init!");
* @tsk is the last thread in the 'cgroup-init' and is exiting.
* Terminate all remaining processes in the namespace and reap them
* before exiting @tsk.
* Note that @tsk (last thread of cgroup-init) may not necessarily
* be the child-reaper (i.e main thread of cgroup-init) of the
* namespace i.e the child_reaper may have already exited.
* Even after a child_reaper exits, we let it inherit orphaned children,
* because, pid_ns->child_reaper remains valid as long as there is
* at least one living sub-thread in the cgroup init.
* This living sub-thread of the cgroup-init will be notified when
* a child inherited by the 'child-reaper' exits (do_notify_parent()
* uses __group_send_sig_info()). Further, when reaping child processes,
* do_wait() iterates over children of all living sub threads.
* i.e even though 'child_reaper' thread is listed as the parent of the
* orphaned children, any living sub-thread in the cgroup-init can
* perform the role of the child_reaper.
NORET_TYPE void do_exit(long code)
struct task_struct *tsk = current;
group_dead = atomic_dec_and_test(&tsk->signal->live);
if (group_dead) {
rc = sys_wait4(-1, NULL, __WALL, NULL);
} while (rc != -ECHILD);
* We can not clear ->child_reaper or leave it alone.
* There may by stealth EXIT_DEAD tasks on ->children,
* forget_original_parent() must move them somewhere.
pid_ns->child_reaper = init_pid_ns.child_reaper;
