Practical Haskell: ByteString, Text, and Strict Data
String — the type every earlier chapter quietly used — turns out to be one of Haskell's least efficient representations for real text. Text, ByteString, and the strict-by-default tools around them are what production code reaches for instead, and this chapter explains exactly when and why.
The problem with String
Every earlier chapter used String freely, because for teaching, its simplicity is exactly the point: String really is just [Char], an ordinary list, with every trick from Chapter 6 onward — map, filter, pattern matching on (x:xs) — applying to it directly. That simplicity has a real cost the moment text stops being a toy example.
Figure: String is a linked list of individually boxed characters — every single letter is its own heap allocation, chained together with cons cells that are also separate heap allocations. Text and ByteString are single packed arrays, however long the content, with no per-character overhead at all.
A five-character String needs ten separate heap objects — five cons cells for the list spine, five boxed Chars for the letters — scattered wherever the garbage collector happened to put them, each one requiring its own pointer to be followed to read the next letter. Data.Text and Data.ByteString store the same five characters in one contiguous block of memory: one allocation, no pointer-chasing, and — for Text specifically — genuinely correct Unicode handling that String’s “just a list of Char” model doesn’t get for free either.
Text: strings, done properly
Data.Text is the type to reach for whenever String was standing in for “human-readable text” — names, messages, file contents meant to be read, anything Unicode-aware.
import qualified Data.Text as T
import qualified Data.Text.IO as TIO
greeting :: T.Text
greeting = T.pack "Hello, world!" -- String -> Text
shout :: T.Text -> T.Text
shout = T.toUpper
main :: IO ()
main = TIO.putStrLn (shout greeting) -- Text-aware IO, not the String-based putStrLn
T.pack/T.unpack convert between String and Text at the boundary; most Text-using code stays in Text throughout, only converting at the edges (parsing command-line arguments, formatting a final String for a library that demands one). OverloadedStrings (covered properly in the next chapter’s extension field guide) lets string literals like "Hello, world!" be Text directly, skipping T.pack almost everywhere.
Text’s Unicode-correctness is not a minor detail String merely handles more slowly — String’s Char-per-element model genuinely gets some Unicode operations wrong by naive implementations (like counting “characters” when a single visible glyph is made of multiple combining code points), the exact kind of bug Text’s library functions are written specifically to avoid.
ByteString: bytes, not characters
Data.ByteString packs the same way Text does, but for raw bytes (Word8), with no notion of “character” or encoding at all. Reach for it when the data genuinely isn’t text — binary file formats, network protocol payloads, image data — or when it’s ASCII-only text where Unicode correctness would be pure overhead.
import qualified Data.ByteString as BS
import qualified Data.ByteString.Char8 as BSC -- ASCII-only convenience functions
readHeader :: FilePath -> IO BS.ByteString
readHeader path = BS.take 4 <$> BS.readFile path -- first 4 bytes, e.g. a magic number
isGzip :: BS.ByteString -> Bool
isGzip bs = BS.take 2 bs == BS.pack [0x1f, 0x8b] -- gzip's magic bytes
ByteString comes in two flavors that are easy to reach for interchangeably without meaning to: Data.ByteString (strict — the whole thing lives in memory as one block, read eagerly) and Data.ByteString.Lazy (built from a list of strict chunks, read incrementally). Lazy ByteString is the right choice for streaming a large file without holding it all in memory at once — precisely the same “generate only as much as gets demanded” idea Chapter 6 built primes on — but reaching for it out of habit on data that’s going to be fully consumed anyway just adds chunk-management overhead for no benefit. Strict ByteString is the sensible default; reach for lazy specifically when streaming genuinely matters.
Strict data: the tuple trap, generalized
Chapter 6’s “tuple trap” showed that foldl' alone doesn’t stop a lazy field inside an otherwise-forced structure from leaking. The general-purpose fix — data types whose fields are strict by construction, so this class of leak becomes structurally impossible — comes from two directions:
{-# LANGUAGE BangPatterns #-}
-- hand-rolled strict fields, one type at a time (Chapter 6's fix)
data Pair a b = Pair !a !b
-- Data.Strict provides ready-made strict versions of the usual suspects
import qualified Data.Strict.Maybe as SMaybe
import qualified Data.Strict.Tuple as STuple
-- SMaybe.Maybe and STuple.Pair guarantee their contents are never a lazy thunk
strict, the package Data.Strict.Maybe and Data.Strict.Tuple come from, exists so nobody has to hand-write Pair a b = Pair !a !b — style strict variants of Maybe, Either, and tuples for every accumulator in every hot loop. Two GHC extensions go a step further and change the default rather than requiring bangs everywhere by hand:
{-# LANGUAGE StrictData #-}
-- every field of every data type declared in this module is strict by default,
-- as if every field had a ! in front of it, without writing the bangs
data Stats = Stats { total :: Int, count :: Int } -- both fields strict, no ! needed
StrictData flips the default for fields; Strict (a separate, more aggressive extension) flips the default for ordinary let/where bindings too, module-wide. Both exist because the standard advice — “add bangs where you find a leak” — requires first finding the leak, and a module dealing entirely in small, hot-path accumulators is often better served by strict-by-default with occasional ~ (lazy annotation) for the rare field that genuinely benefits from staying lazy, rather than the reverse.
None of this changes anything about purity (Chapter 8) — a StrictData field is still computed from its inputs with no side effects, evaluated eagerly rather than on demand. Strictness and purity are independent axes: Chapter 6 already showed ST’s local mutability staying pure despite genuine in-place writes, and strict fields are the same story one level simpler — when something is computed changes, what it’s allowed to do while computing never does.
IORef: everyday mutable state in IO
Purity’s ST monad (Chapter 8) gave mutation a scoped, provably-pure escape hatch — genuine in-place writes, but only ever visible inside one runST computation, never leaking out. Ordinary IO code frequently wants something simpler: a single mutable cell, alive for as long as the program needs it, with no runST boundary to manage. IORef is exactly that:
import Data.IORef
main :: IO ()
main = do
counter <- newIORef (0 :: Int)
writeIORef counter 5
modifyIORef counter (+1)
final <- readIORef counter
print final -- 6
newIORef, readIORef, writeIORef, and modifyIORef are the whole interface — genuinely just a mutable box, with no purity trick hiding underneath the way ST’s state token has one. That’s precisely why IORef lives in IO rather than getting ST’s special treatment: reading and writing it is a real side effect, visible to anything else sharing the same reference, and the type honestly says so.
modifyIORef is lazy in the new value by default — modifyIORef counter (+1) builds a thunk old + 1 rather than computing it immediately, which is exactly Chapter 6’s tuple-trap thunk-buildup risk, now happening to a long-lived mutable cell instead of a fold accumulator. modifyIORef' (with a trailing prime, matching foldl'’s own naming convention) forces the new value immediately instead, and is the safer default for any IORef updated frequently in a loop.
Handling exceptions without losing your footing
Maybe and Either (Chapters 5 and 11) handle expected failure — a lookup that might miss, a parse that might fail — as an honest part of a function’s own type. Real IO code also has to deal with unexpected failure: a file that doesn’t exist, a network call that times out, division by zero. Control.Exception is where that lives:
import Control.Exception
safeDiv :: Int -> Int -> IO Int
safeDiv x y = evaluate (x `div` y) `catch` \e -> do
putStrLn ("Caught: " ++ show (e :: ArithException))
return 0
ghci> safeDiv 10 2
5
ghci> safeDiv 10 0
Caught: divide by zero
0
catch takes an IO action and a handler, running the handler if the action throws an exception matching the handler’s argument type — ArithException here, so catch only intercepts arithmetic-related exceptions, letting anything else propagate normally.
The evaluate above is not decoration — removing it breaks the example in a genuinely surprising way. return (x \div` y)builds anIO Intaround an *unevaluated* thunk; the division itself doesn't happen until something later actually demands theInt, which might well be *after* catch's handler has already finished running and moved on. evaluate :: a -> IO aforces its argument to WHNF immediately, inside theIOaction, which is precisely what makes sure the exception fires *whilecatch` is still watching* — a sharp, easy-to-miss interaction between Chapter 6’s laziness and exception handling that catches even experienced Haskell programmers occasionally.
A second, equally common need is guaranteeing cleanup happens even when something goes wrong partway through — closing a file handle whether reading it succeeded or crashed halfway:
import System.IO
withFileSafely :: FilePath -> (Handle -> IO a) -> IO a
withFileSafely path = bracket (openFile path ReadMode) hClose
bracket acquire release use runs acquire, then use on the result, and guarantees release runs afterward — whether use finished normally, returned early, or threw an exception. This is the same resource-safety instinct C++‘s RAII or Python’s with blocks reach for, expressed as an ordinary higher-order function rather than new syntax.
catch and bracket both operate on ordinary IO values, with no special language support — they’re implemented in terms of the RTS’s own exception-delivery mechanism (The Runtime System and Performance Tooling), exposed as regular functions. This is the same pattern Parser Combinators and EDSL built entire chapters around: powerful behavior, implemented as a library, rather than baked into the compiler.
Choosing, in practice
| Situation | Reach for |
|---|---|
| Human-readable text, Unicode matters | Text |
| Binary data, network payloads, ASCII-only bytes | ByteString (strict) |
| Streaming a large file without loading it all | ByteString.Lazy or Text.Lazy |
| An accumulator in a hot loop (tuple, record) | strict fields (!, StrictData, or strict’s ready-made types) |
| Everything else — small scripts, teaching, glue code | String is still completely fine |
Real production Haskell overwhelmingly prefers Text and ByteString over String the moment code touches genuine user input, file I/O, or network traffic — web frameworks like Yesod and Servant, JSON libraries like aeson, and HTTP clients like http-client all default to Text/ByteString in their core APIs, precisely because the memory-layout difference this chapter’s figure shows stops being academic at real traffic volumes. String remains extremely common in Haskell tooling and small scripts, where its simplicity outweighs the cost at the scale involved.
None of Text, ByteString, or strict data change what a program computes — every example in this chapter could, in principle, be rewritten with plain String and lazy tuples and produce the identical answer. What changes is how much the machine has to do, and how much memory it holds onto, to get there — the exact gap this book has been calling “clarity first, performance once it’s earned its place” since the very first Pearls chapter.