Created by a human, with a brain badgeA badge with a character next to the text "Web 14," indicating that the site may contain slightly offensive materialDo What The Fuck You Want To Public License badgeD D Race Network badgemi toki e Toki Ponamade with MY OWN TWO PAWS badgeNo cookies badgeNo tracking or analytics badgeMade with server-side scripting badgeHosted on home internet badge


The Worst Thing in Unix

I have lost days of my life to this behavior, both in supremely frustrating whole blocks and in frequency. It manifests at random in my everyday, and often I find it quite difficult to realize that it is happening even now that I am well aware of it. Once you do realize that it is the root cause of all of your problems, it is horrible to debug and you will hate any fix you implement.

Join Me In Horror

Open up a shell and run the following:
time sh -c '{ printf "\n"; sleep 1; printf "\n"; sleep 4; printf "I lived!\n" >&2; } | read'


The first newline is read by read, the second triggers a SIGPIPE, and we never get a joyful indication of life. Great!
Now try:
time sh -c '{ sleep 1; printf "\n\n"; sleep 4; printf "I lived!\n" >&2; } | read'


Oh no. Ah, but both lines were sent in one write call, so surely it makes sense to return how much was written instead of possibly -1 with errno=EPIPE depending on the CPU scheduling.
Let's try something with more definable behavior - surely either the second call gets SIGPIPEd or we get one late as the pipe buffer "flushes," right? Right?
time sh -c '{ printf "\n"; sleep 1; printf "\n"; sleep 5; printf "I lived!\n" >&2; } | { sleep 2; read; }'


Oh dear. Oh dear.

For those not near a terminal, the second two live for 5s instead of 1s and print "I lived!"

Consequences

Your writer lives possibly forever! Parents waiting on hanging children! ({ printf "\n"; sleep inf; } | sleep 1) Leaked PIDs and resources! (bash -c 'read < <(printf "\n"; sleep inf)', check ps)

You cannot ever rely on SIGPIPE to end a process unless you know that it will regularly produce output that is not filtered out until the very end of your pipe chain.

In many cases, that is not an easy task. Your only workarounds in Bash if you cannot guarantee a SIGPIPE chain are coprocesses (of which you can only. . . legally. . . have one, currently) or manually doing things with FIFOs and manual process management (which gets messy quickly).
Let me give you a case study. Let's talk about the occurance that caused me to write this all up.

Explanation

Pipes

XKCD 2501 and all that, so here is what is going on under the hood. (Its alt text may be particularly apt, as I have never heard of anyone else having an issue with this behavior: "How could anyone consider themselves a well-rounded adult without a basic understanding of silicate geochemistry? Silicates are everywhere! It's hard to throw a rock without throwing one!")

The most common way for two programs to communicate in Unix systems is via a pipe. | is a shell operator that pipes the output of its left command to the input of its right. A pipe appears to each of the programs as a file, but what is written to the write end can (by and large) only be read in sequence and once from the read end. If the reader tries to read more data than is available, it will generally be (transparently, to it,) stalled until more is written.

The sticky bit is the equivalent behavior from the write end. Pausing every time something is written in order to go hand it to the reader would be terribly slow, and with our wonderful modern technology we can do more than one thing at once. Instead there is a pipe buffer, a certain amount which can be written without the writer being stalled. If both processes were to produce and consume at the same pace, neither would wait and we could proceed as fast as possible.
(In modern Linux kernel versions the default pipe buffer is 16 memory pages, usually 64KiB which is around eleven thousand words.)

The Problem

If a program tries to write to a pipe whose other end has exited, it gets a signal called SIGPIPE. SIGPIPE makes it aware of that fact, and it can then proceed accordingly. Most programs maintain the default behavior, which is to immediately exit. This, in my opinion horribly erroneous, behavior occurs when the pipe buffer is written to by the writer (and flushed!) and then the reader reads either none or some of that content. Until the writer writes again, after the reader has exited, it does not get SIGPIPEd.

Case Study

At my work we use virtio-snd to pass audio streams to and from virtual machines. If you start using the virtual devices in the VM too early things will lock up in the host, and if you start using them after that but still too early things will lock up in the guest. That all sounds like a Later Problem, but for now, I need a script to wait until the devices are available in PipeWire (which seems to be a variable amount of time) and then sleep for a few seconds.

That sounds easy enough - we can use pw-mon, filter for entries matching our node.names and such, and exit once we see them all! Ah ah ah, but SIGPIPE is not our friend. You see, if there are no extra entries put down our pipeline after we see them all, or if those extra entries arrive too early, the head of our pipeline (and thus our script) will never exit. In fact, in this case that is guaranteed to happen, because the whole point is to not cause further events with those devices until the script exits!

Let's examine our two workarounds:
  • Coprocesses. Of note, we have to filter the data. pw-mon has a lot of output on our system, and filtering in Bash was very slow. If we wanted to use grep or sed outside of the coprocess, we would have to either use process substitution or copy the coprocess PID for killing when finished (coprocess details are not available in subprocesses which a pipe from grep/sed would be, GNU doesn't trust you) - but it would still work just fine. In my case there is no problem with piping to the filter inside of the coprocess, and I think that is an ideal and clean solution. Unfortunately I only had ash, so it was not an option.
  • FIFOs and manual process management. I generally think this should be avoided, it is too easy to make zombie processes waiting on nonexistent FIFOs even if only during development. Regardless, I would only need one and I realize that it is better in this situation after sleeping on it. I may rewrite to this approach next week but am still conflicted over the grep (we'll get there) and I may just stop wasting time on this stupid thing now that it is working.
I went with neither, as I have also sometimes done in situations where I have less control over the system or if something anywhere in the chain could cause the exit of the pipeline. Here is what a SIGPIPE chain for this problem looks like:
( pw-mon & A & ) | ( B | grep & ) | ( C & ) | D

A prints a newline every second, and kills pw-mon if it is SIGPIPEd. B forwards those newlines directly to C, and everything else to grep for filtering (so that grep's internal buffer doesn't have to fill before SIGPIPE is possible). C forwards the newlines and does some requisite logic, and D is the core logic. There is lots of file descriptor redirecting in there. Everything is doubly backgrounded (disowned) so that the script exits immediately when D does, and everything gets cleaned up by SIGPIPE a second later.

Further Buffering

Even if I am just off in my own world, I expect you may have run into this behavior due to how it is exacerbated by userspace buffering. See here (archive) for a good page on how this is otherwise annoying and may cause identical issues, but programs printing things in large blocks makes it very common for them to delay being SIGPIPEd or not be SIGPIPEd at all in the same way as my first bad example.

I am lucky in that the grep in my case above does not block! --line-buffered is not available, but thankfully despite there not being enough data to avoid the above issue it somehow seems to work. I am not sure why. There is a timeout which would cause B to exit, causing grep to exit from the front, but it flushes before then and the timeout is not triggered. Especially given that additional possible complication, I am so far beyond the point where I would like to use another language and do everything myself straight from pw-mon, even with worse pipeline and string handling, if I had a good way of deploying something written in one. Maybe I could use sed and store some state in the hold space so that everything is done in one sed process in a worst-case scenario? That is horrible and noone but me will be able to read it, though. All I wanted to do was wait for a few specific lines from a process, THIS IS WHAT SHELLS ARE MADE FOR! pw-mon exists only to provide lines for me to filter on, that should not be a hard thing to express.

Possible Solutions

I do believe this is solvable, with either a kernel-space or user-space implementation:
  • In the kernel-space, the length of the last write() could be recorded, and if the reader exits before reading into the last <that many> bytes then the writer should be SIGPIPEd. There would be complications when stacked with additional userspace buffering, but with some cooperation that can be addressed. Doing so would break the (AFAIK) current contract of only being able to receive a SIGPIPE while actively doing things with the file descriptor in question.
  • In the user-space, I believe it is possible to monitor your FDs. It doesn't make sense for all programs to do so, but I am working on a little passthrough helper which can watch and eagerly SIGPIPE its neighbors. I will update this page once I finish or fail at that.