Category Theory Key idea: A domain-specific language doesn't need its own compiler — if the host language is expressive enough, the DSL can just be a very well-designed library

Embedded Domain-Specific Languages

Parser Combinators and Optics were already embedded domain-specific languages, without ever being named as such. This chapter names the pattern properly, explains exactly why Haskell is unusually suited to it, and surveys the real, production libraries — Diagrams, Servant, Shake, Esqueleto — that lean on the same idea.

What makes a DSL “embedded”

A domain-specific language is any language built for one narrow job rather than general-purpose programming — SQL for queries, regular expressions for pattern matching, Make’s syntax for build dependencies. The traditional way to build one is to build an entire language: a lexer to tokenize it, a parser to turn tokens into a syntax tree, a type checker if it has types, and an interpreter or compiler to actually run it.

An embedded domain-specific language (EDSL) skips nearly all of that. Instead of inventing new syntax, an EDSL is a library, written in an existing “host” language, whose functions and operators are deliberately designed to read like a specialized language even though they’re really just ordinary function calls.

Comparison of a standalone DSL's four custom-built pipeline stages versus an embedded DSL reusing three of the host language's own stages

Figure: Three of the four stages are simply GHC’s own, already correct, already fast, with editor support already built in — only the last stage, deciding what the DSL’s combinators actually mean, is genuinely new code.

Parser Combinators, several chapters back, was already exactly this without ever using the word: charP, <*>, and <|> aren’t Haskell syntax extensions — they’re ordinary functions and operators, and pointP = Point <$> intP <* charP ',' <*> intP reads like a small, purpose-built grammar description precisely because it is one, expressed entirely in vanilla Haskell. Optics did the same thing for navigating nested data. Neither chapter needed a single line of custom parsing infrastructure, because the “DSL” was never a separate language to begin with — just a library used in a particular style.

Why Haskell specifically is good at this

Plenty of languages can define functions; embedding a genuinely readable domain language takes more than that. A handful of Haskell features work together to make it unusually natural:

Custom infix operators. <*>, <$>, <>, and <|> are not built into Haskell’s grammar as special cases — they’re ordinary functions with symbolic names, and any library can define its own the same way, with a chosen precedence and associativity. A domain vocabulary gets to look like operators, not just function calls buried in parentheses.

Typeclasses as a source of “operator overloading.” The same <> means list concatenation, numeric addition (via Sum), or HTML-tag concatenation (this chapter, shortly) depending on which Monoid instance is in scope — one familiar symbol, reused meaningfully across completely different domains, resolved automatically by the type checker rather than requiring a different symbol for every context.

Do-notation, generalized. Chapter 11 already established that do-blocks work for any Monad, not just IO — which means defining a custom Monad instance for a domain type gives that domain an entire imperative-looking mini-syntax, for free, with zero new parser code. A domain-specific “sequence of steps” language is often just a custom Monad instance away.

Laziness. An EDSL describing an infinite or self-referential structure — a game board, a stream of events, a recursively-defined shape — can be written exactly the way the domain expert would draw it, with Haskell’s evaluation model quietly ensuring only the needed parts are ever actually built (Chapter 6).

A type system expressive enough to catch domain errors. Type Families’ Servant example encoded an entire REST API’s shape as a type — meaning a malformed API description isn’t a runtime bug to discover later, it’s a compile error, caught before the program runs at all.

λA Category-Theoretic View

This is the same “why Applicative was worth learning once, abstractly” argument Parser Combinators made concrete: an EDSL’s combinators are frequently just ordinary Functor, Applicative, Monad, or Monoid instances for a domain-specific type, which means the interface — <*>, <>, do-notation — was already generic, already law-abiding, and already familiar, long before this particular domain existed. Embedding a new DSL rarely means inventing new abstractions; it usually means finding which abstraction this book already covers actually fits the domain.

Building one from scratch: an HTML EDSL

Monoids, two chapters back, gave every type with an associative combining operation a shared vocabulary — <> to combine, mempty as the identity. HTML markup turns out to fit that shape exactly: two pieces of markup concatenate into one, and there’s an obvious “nothing” (empty markup).

newtype Html = Html String

instance Semigroup Html where
  Html a <> Html b = Html (a ++ b)

instance Monoid Html where
  mempty = Html ""

tag :: String -> Html -> Html
tag name (Html inner) = Html ("<" ++ name ++ ">" ++ inner ++ "</" ++ name ++ ">")

text :: String -> Html
text = Html

toString :: Html -> String
toString (Html s) = s

With that foundation, the actual domain vocabulary is nothing more than a handful of one-line combinators, each just tag partially applied to a specific tag name:

html_, body_, h1_, p_ :: Html -> Html
html_ = tag "html"
body_ = tag "body"
h1_   = tag "h1"
p_    = tag "p"

page :: Html
page = html_ (
    body_ (
      h1_ (text "Welcome") <>
      p_ (text "This is a paragraph.")
    )
  )

ghci> toString page
"<html><body><h1>Welcome</h1><p>This is a paragraph.</p></body></html>"

page reads like a small, declarative description of a document’s structure — because that’s exactly what it is, built entirely from tag, text, and the <> this book already spent a full chapter justifying. No new parser was written; GHC’s own parser handled every line, and GHC’s own type checker ensures page really is an Html value before anything runs.

⚠Common Pitfall

This toy Html type has an obvious hole: nothing stops tag "h1" from being applied where a <title> was expected, or an attribute value from containing unescaped < and > characters that corrupt the output. Real HTML EDSLs (below) close exactly these gaps with more elaborate types — separate types for “things valid inside <body>” versus “things valid inside <head>,” and automatic escaping built into text itself — but the underlying combinator shape, tags built from smaller tags via a Monoid-flavored operator, stays the same.

A second shape: describing a program, then interpreting it separately

The HTML example combines values directly into a result. A different, equally common EDSL shape keeps the description of what should happen entirely separate from what actually happens when it runs — letting the same description be interpreted more than one way.

data Dir = North | East | South | West deriving (Show, Eq, Enum, Bounded)

turnRight, turnLeft :: Dir -> Dir
turnRight West = North
turnRight d    = succ d
turnLeft North = West
turnLeft d     = pred d

data Cmd = Forward Int | TurnLeft | TurnRight deriving Show
type Program = [Cmd]

square :: Int -> Program
square n = concat (replicate 4 [Forward n, TurnRight])

square builds a Program — an ordinary list of Cmd values, plain data, doing nothing on its own. What that data means is decided entirely separately, by whichever interpreter is run over it:

-- interpreter 1: render the program as readable text
describe :: Program -> String
describe = unlines . map show

-- interpreter 2: actually simulate it
data Robot = Robot { pos :: (Int, Int), dir :: Dir } deriving Show

run :: Program -> Robot
run = foldl step (Robot (0, 0) North)
  where
    step (Robot (x, y) d) (Forward n) = Robot (move d) d
      where move North = (x, y + n)
            move South = (x, y - n)
            move East  = (x + n, y)
            move West  = (x - n, y)
    step (Robot p d) TurnLeft  = Robot p (turnLeft d)
    step (Robot p d) TurnRight = Robot p (turnRight d)

ghci> run (square 5)
Robot {pos = (0,0), dir = North}

The same four-command square 5 program can be pretty-printed by describe, actually executed by run, or — without changing a single line of square itself — handed to a third interpreter that renders it as an SVG path, logs every command for debugging, or checks it never leaves some bounded area. This separation, description as inert data with interpretation supplied afterward, is one of the deepest reasons EDSLs are worth the design effort: the domain vocabulary gets defined exactly once, and every new way of using that vocabulary is a new interpreter, not a rewrite.

★Cool Fact

This “build a data structure describing what to do, then write separate functions that decide what it means” pattern has a name when taken to its logical extreme: free monads, from Effects Beyond IO’s neighborhood of ideas, generalize exactly this description-then-interpretation split into a reusable, principled structure — letting an entire domain-specific effect (not just a list of commands) be described once and interpreted however many different ways a program needs, including a real one for production and a fake one for tests.

A survey of real, production Haskell EDSLs

Every idea above shows up, considerably more developed, in real, widely-used libraries:

  • Diagrams — a declarative vector-graphics EDSL where pictures are built by combining smaller pictures with combinators like (|||) (place beside) and (===) (place above), directly echoing this chapter’s HTML-combinator shape, but for shapes and layout instead of markup.
  • Servant — Type Families’ own example, worth remembering here specifically as an EDSL: an entire REST API’s shape, encoded as a type, checked by GHC’s own type checker with zero custom tooling written.
  • Shake — a build system, in the spirit of Make, where dependency rules are ordinary Haskell functions and values rather than a separate Makefile syntax, gaining real programming-language power (loops, abstraction, real data structures) that Make’s own restricted syntax can’t offer.
  • Esqueleto — type-safe SQL query construction, where a malformed join or a column reference to a table that isn’t in scope is a compile error, not a runtime SQL exception discovered in production.
  • Hspec — a testing EDSL built almost entirely from ordinary function calls and do-notation, producing human-readable test specifications (describe "reverse" $ it "is its own inverse" $ ...) that still compile and run as completely ordinary Haskell.
  • Blaze and Lucid — the real, production versions of this chapter’s toy Html type, with proper escaping, attribute handling, and a considerably richer combinator vocabulary, but recognizably the same Monoid-flavored design at heart.
In the Wild

The EDSL style is common enough in Haskell that reaching for a combinator library, rather than writing a bespoke parser and interpreter, is often the default instinct for a new domain-specific need — precisely because so much of the hard work (a correct parser, a sound type checker, mature tooling) comes free the moment the “language” is just well-designed Haskell. Domains as different as build systems, SQL, vector graphics, and web APIs have all converged on the same underlying technique independently, which is a good sign the technique itself is the real discovery, not any one of its applications.

Every chapter in this book’s “Further Peaks” section has, in its own way, been building toward this one: Optics turned “navigate a nested structure” into a combinator vocabulary; Parser Combinators turned “describe a grammar” into one; Monoids gave this chapter’s HTML example its combining operator for free. An EDSL is rarely a new invention — it’s usually this book’s own toolbox, aimed at one more domain.