Ruby

Escaping Characters in Ruby Strings and Regular Expressions

Escaping Characters in Ruby Strings and Regular Expressions

In Ruby, a backslash escapes the character after it. In a double-quoted string, \" inserts a quote, \\ a backslash, \n a newline, and \#{} blocks interpolation. A single-quoted string recognizes only \' and \\; every other backslash stays literal. When escaping gets noisy, switch to %q()/%Q() literals or a heredoc.

That is the whole rule; the rest is detail: which sequences double quotes recognize, how heredocs and percent literals cut the backslashes down, why inspect shows escapes you never typed, how to escape text that reaches a regular expression at runtime, and what Ruby 3.4 changed about string literals. Every example on this page was regenerated on Ruby 3.4.10; the complete escape table lives in Ruby’s literal syntax documentation. If you landed here with an error message, jump to the error reference.

Escaping Quotes

A string needs a backslash before one character only: the quote that opened it. Inside double quotes, write a double quote as \"; inside single quotes, write a single quote as \':

Ruby
"Hello \"world\"!" # => "Hello \"world\"!"
'Hello \'world\'!' # => "Hello 'world'!"

The other quote character needs nothing, so the shortest fix for a quote inside a string is to open the string with the other kind:

Ruby
"It's a string" # => "It's a string"
'Say "hi"'      # => "Say \"hi\""

The display is what confuses people. The # => comments show what inspect returns, and inspect renders every string between double quotes. It puts a backslash before a double quote, before a backslash, and before control characters such as a newline; a single quote is printed as it is:

Ruby
'"' # => "\""
"'" # => "'"

So the backslash in "Say \"hi\"" exists in the display, not in the string. The string holds a plain quote character.

Percent Literals

Percent literals move the escaping problem from quotes to delimiters. %() and %Q() behave like a double-quoted string, %q() like a single-quoted one, and none of them care about quote characters inside:

Ruby
%(Hello "world"!) # => "Hello \"world\"!"

The delimiter pair is what you escape now, and only when it is unbalanced. A balanced pair nests without help:

Ruby
%(Hello (world)!) # => "Hello (world)!"
%(Hello world\)!) # => "Hello world)!"

Escape sequences and interpolation follow the single-versus-double rule:

Ruby
name = "world"
 
%(Hello\n "#{name}"!)  # => "Hello\n \"world\"!"
%Q(Hello\n "#{name}"!) # => "Hello\n \"world\"!"
%q(Hello\n "#{name}"!) # => "Hello\\n \"\#{name}\"!"

Any non-alphanumeric character can be the delimiter, which is how you pick one that does not appear in the text. Even quote characters work, without a warning under -w:

Ruby
%[foo] # => "foo"
%{foo} # => "foo"
%<foo> # => "foo"
%|foo| # => "foo"
%-foo- # => "foo"
%!foo! # => "foo"
%"foo" # => "foo"
%'foo' # => "foo"

The word-array form %w[] is the one place a backslash before a space matters: it keeps the two halves in one element.

Ruby
%w[a\ b c] # => ["a b", "c"]

The full list of percent literals is in Ruby’s literals documentation.

Heredocs

A heredoc holds several lines of text with no quote escaping at all. The squiggly form <<~EOS strips the common leading indentation and keeps escape sequences and interpolation on. Quote the terminator in single quotes to turn both off, the way a single-quoted string does:

Ruby
name = "world"
 
text = <<~EOS
  Hello #{name}
    indented \t tab
  "quotes" and 'quotes' need no escaping
EOS
text # => "Hello world\n  indented \t tab\n\"quotes\" and 'quotes' need no escaping\n"
 
raw = <<~'EOS'
  Hello #{name}
  no \t escapes either
EOS
raw # => "Hello \#{name}\nno \\t escapes either\n"

<<~"EOS" means the same as <<~EOS. The older <<-EOS form only lets you indent the terminator; the body keeps its indentation. Two more details: a backslash at the end of a line inside a squiggly heredoc joins it to the next line, and a method call chains off the opener:

Ruby
dash = <<-EOS
  Hello world
    keeps indentation
  EOS
dash # => "  Hello world\n    keeps indentation\n"
 
joined = <<~EOS
  Line with a trailing backslash \
  joins the next line
EOS
joined # => "Line with a trailing backslash joins the next line\n"
 
shout = <<~EOS.upcase
  hello
EOS
shout # => "HELLO\n"

Here document literals in the Ruby docs cover the remaining forms.

Escape Sequences

A backslash followed by a letter can also stand for a character you cannot type into a literal. \n is a newline, and the pair is called an escape sequence:

Ruby
"Hello\nworld" # => "Hello\nworld"

The returned string is displayed with the sequence intact, because inspect escapes control characters. Print it and the newline is real:

Ruby
puts "Hello\nworld"
text
Hello
world

Single quotes do not process escape sequences; '\n' is a backslash followed by an n. In a double-quoted string, you write that pair as \\n:

Ruby
'\n'            # => "\\n"
"Hello\\nworld" # => "Hello\\nworld"

The Escape Sequence Table

Every sequence a double-quoted string (or %Q(), or a heredoc) recognizes, with what it produces on Ruby 3.4:

SequenceMeaningExample
\nnewline, byte 10"\n".bytes → [10]
\ttab, byte 9"\t".bytes → [9]
\sspace, byte 32"\s" == " " → true
\0NUL, byte 0"\0" → "\u0000"
\eescape, byte 27"\e[31mred\e[0m" colors terminal output
\abell, byte 7"\a".bytes → [7]
\bbackspace, byte 8"\b".bytes → [8]
\rcarriage return, byte 13"\r".bytes → [13]
\fform feed, byte 12"\f".bytes → [12]
\vvertical tab, byte 11"\v".bytes → [11]
\\one backslash"\\".length → 1
\"double quote"\"" → "\""
\'single quote; the backslash is dropped"\'" → "'"
\#hash sign; only needed before {, $, or @"\#" → "#"
\uXXXX, \u{…}Unicode code point; braces take one or more, separated by spaces"\u00e9" → "é", "\u{48 49 21}" → "HI!", "\u{1F600}" → "😀"
\xNNone byte, hexadecimal"\x41" → "A", "\xC3\xA9" == "é" → true
\NNNone byte, octal"\101" → "A", "\1" → "\u0001"
\cx, \C-xcontrol character"\C-a".bytes → [1]
\M-xmeta character, high bit set; one byte that is not valid UTF-8"\M-a".bytes → [225]
\M-\C-xmeta and control"\M-\C-a".bytes → [129]
backslash, newlineline continuation inside the literal; both characters vanishsee the following example
Ruby
"a\
b" # => "ab"

A backslash before any character not in the table is dropped without a warning, even under -w: "\q" is "q". Regex literals are stricter about that, as the regular expression section shows. Ruby’s escape sequence table is the canonical list.

Single Quotes vs Double Quotes

Single quotes recognize exactly two escapes: \\ for a backslash and \' for a single quote. Every other backslash stays in the string, including one before a double quote, a space, or a newline:

Ruby
'\\'   # => "\\"
'\''   # => "'"
'\"'   # => "\\\""
'a\ b' # => "a\\ b"
'a\
b'     # => "a\\\nb"

Both kinds produce ordinary String objects; the difference is only in how Ruby reads the literal. Use double quotes when you want the sequences interpreted, single quotes when you want to show them:

Ruby
puts "Line 1\nLine 2"
puts 'Using a \n we can indicate a newline.'
text
Line 1
Line 2
Using a \n we can indicate a newline.

That rule answers the Windows path question. In double quotes, every backslash in the path needs doubling; in single quotes, none of them do:

Ruby
puts "C:\\Users\\tom"
puts 'C:\Users\tom'
text
C:\Users\tom
C:\Users\tom

Seeing Escapes with inspect, p, puts, and dump

Two views of one string cause most of the confusion in this topic. puts and print write the interpreted characters; p and inspect show the literal form, with escapes added so you could paste it back into a program:

Ruby
s = "Tab\there \"quoted\" 'single'"
 
puts s
p s
text
Tab	here "quoted" 'single'
"Tab\there \"quoted\" 'single'"

s.inspect returns that second line as a String, and p s does the same as puts s.inspect. String#dump goes one step further and escapes every non-ASCII character too, which makes its output safe for a file that cannot hold UTF-8; undump reverses it exactly:

Ruby
"café\n".inspect # => "\"café\\n\""
"café\n".dump    # => "\"caf\\u00E9\\n\""
"café\n".dump.undump == "café\n" # => true

Bytes that do not form valid UTF-8 show up in inspect as \x escapes, and valid_encoding? tells you the string is broken before a method raises on it:

Ruby
"\xff"                 # => "\xFF"
"\xff".valid_encoding? # => false

The Ruby docs describe String#dump and String#undump in full.

Escaping Interpolation

A single-quoted string holds the characters #{name} as they are, and the display adds a backslash to say so:

Ruby
name = "world"
 
"Hello #{name}" # => "Hello world"
'Hello #{name}' # => "Hello \#{name}"

To get the same literal text inside double quotes, escape the hash sign:

Ruby
"Hello \#{name}" # => "Hello \#{name}"
puts "Hello \#{name}"
text
Hello #{name}

Two shorthand forms exist that many people never see. A global or instance variable interpolates with no braces at all, and the same backslash turns those off. A hash sign followed by anything else is plain text:

Ruby
$name = "global"
@name = "ivar"
 
"Hello #$name"   # => "Hello global"
"Hello #@name"   # => "Hello ivar"
"Hello \#$name"  # => "Hello \#$name"
"Hello \#@name"  # => "Hello \#@name"
"Hello #name"    # => "Hello #name"
"Hello # {name}" # => "Hello # {name}"

The backslash in those displays is one you never typed. inspect escapes #{, #$, and #@, and nothing else, so that its output re-parses as the same string:

Ruby
'#$name' # => "\#$name"
'#name'  # => "#name"

When a string is mostly literal hash signs, skip the escaping: %q() and <<~'EOS' never interpolate. Interpolation follows the same single-versus-double rule as escape sequences: double quotes support it, single quotes don’t.

Escaping Characters in Regular Expressions

In a regular expression, many characters mean more than themselves: . matches any character, [] selects from a set, () groups. A backslash gives each its literal meaning:

Ruby
/Hello \[world\]/ # => /Hello \[world\]/
/\[world\]/ =~ "Hello [world]" # => 6

Without the escapes, [world] is a character class, and the first match is the l of Hello, at index 2:

Ruby
/[world]/ =~ "Hello [world]" # => 2

The forward slash that delimits a regex literal needs escaping too, unless you change the delimiter with %r{}. Ruby adds the escapes for you in the display, and the two literals are equal:

Ruby
%r{/world/}        # => /\/world\//
%r{/world/}.source # => "/world/"
%r{/world/} == /\/world\// # => true

One asymmetry with strings: a regex literal warns about an unknown escape. Under ruby -w, /\q/ prints warning: Unknown escape \q is ignored: /\q/, where "\q" stays silent.

Escaping Dynamic Input with Regexp.escape

Text that arrives at runtime, a search term or a file name, cannot be escaped by hand. Interpolate it as it is and every metacharacter in it is live:

Ruby
input = "a.b"
 
"axb".match?(/#{input}/)                # => true
"axb".match?(/#{Regexp.escape(input)}/) # => false
"a.b".match?(/#{Regexp.escape(input)}/) # => true

Regexp.escape (alias Regexp.quote) puts a backslash before every metacharacter, and also turns a space into a backslash-space pair and a tab into \t. Regexp.union escapes its string arguments for you and leaves Regexp arguments alone:

Ruby
Regexp.escape("1+1=2? (yes.)")   # => "1\\+1=2\\?\\ \\(yes\\.\\)"
Regexp.new(Regexp.escape(input)) # => /a\.b/
 
Regexp.union("a.b", "c*d") # => /a\.b|c\*d/
Regexp.union("a.b", /c*d/) # => /a\.b|(?-mix:c*d)/

An unescaped ( or [ in interpolated input does not match wrongly; it raises. That message, end pattern with unmatched parenthesis, has its own entry in the error reference. The Ruby docs cover Regexp.escape and Regexp.union.

Escaping Line Breaks

The backslash also escapes a line break in Ruby code itself. A method call that runs past your line length limit can continue on the next line as long as the previous line ends with a comma, an operator, or an open parenthesis:

Ruby
def ruby_method_with_many_arguments(str, split:, join:)
  str.split(split).join(join)
end
 
ruby_method_with_many_arguments "Hello world", split: " ", join: "\n" # => "Hello\nworld"
 
ruby_method_with_many_arguments "Hello world",
  split: " ",
  join: "\n" # => "Hello\nworld"
 
ruby_method_with_many_arguments(
  "Hello world",
  split: " ",
  join: "\n"
) # => "Hello\nworld"

To put every argument on its own line at the same indentation, without parentheses, end the first line with a backslash:

Ruby
my_string = "Hello world"
 
ruby_method_with_many_arguments \
  my_string,
  split: " ",
  join: "\n" # => "Hello\nworld"

The backslash does not escape the newline character; it tells the parser that the statement continues. The next line therefore has to be something that can continue the expression. A bare 42 on the line after "foo" \ is a syntax error, unexpected integer, expecting end-of-input.

Escaping Line Breaks in String Definitions

The same continuation works inside a string definition. The usual way to split a long string across lines is +:

Ruby
"foo" +
  "bar" # => "foobar"

A backslash at the end of the line gives the same result:

Ruby
"foo" \
  "bar" # => "foobar"

The two are not the same operation. Ruby concatenates adjacent string literals at parse time, whether they sit on one line or are joined with a backslash. The chain can be as long as you like:

Ruby
"foo" "bar" # => "foobar"
"foo" 'bar' # => "foobar"
 
"foo" \
  "bar" \
  "baz" # => "foobarbaz"

Ask Ruby to disassemble both forms and the difference is visible. The + version pushes two string literals and calls +, which allocates a third string at runtime:

Ruby
puts RubyVM::InstructionSequence.compile(<<~'RUBY').disasm
  "foo" +
    "bar"
RUBY
text
== disasm: #<ISeq:<compiled>@<compiled>:1 (1,0)-(2,7)>
0000 putchilledstring                       "foo"                     (   1)[Li]
0002 putchilledstring                       "bar"                     (   2)
0004 opt_plus                               <calldata!mid:+, argc:1, ARGS_SIMPLE>(   1)[CcCr]
0006 leave

The backslash version compiles to a single literal:

Ruby
puts RubyVM::InstructionSequence.compile(<<~'RUBY').disasm
  "foo" \
    "bar"
RUBY
text
== disasm: #<ISeq:<compiled>@<compiled>:1 (1,0)-(2,7)>
0000 putchilledstring                       "foobar"                  (   1)[Li]
0002 leave

One string object instead of three, which is why the backslash form is the better choice for long literals. The Ruby Magic post on how Ruby compiles this example walks through the instruction sequence. The instruction name, putchilledstring, points at a Ruby 3.4 change to string literals, the subject of the next section.

Chilled String Literals in Ruby 3.4

Ruby 3.4 made bare string literals “chilled”: still mutable, not frozen, but marked so the interpreter can warn when a program mutates one. A future version will freeze them. The Ruby 3.4 NEWS puts it this way: “String literals in files without a frozen_string_literal comment now emit a deprecation warning when they are mutated. These warnings can be enabled with -W:deprecated or by setting Warning[:deprecated] = true. To disable this change, you can run Ruby with the --disable-frozen-string-literal command line argument.” Nothing changes in the values:

Ruby
s = "hello"
s.frozen?     # => false
s << " world" # => "hello world"
s.frozen?     # => false

The warning prints only when deprecation warnings are on: -w, -W2, -W:deprecated, RUBYOPT=-W:deprecated, or Warning[:deprecated] = true in code. It does not print by default, not under -W1, not in a default irb session (Warning[:deprecated] is false there), and not with --debug-frozen-string-literal on its own. Save these three lines as chilled.rb:

Ruby
s = "hello"
s << " world"
p s
text
$ ruby -W:deprecated chilled.rb
chilled.rb:2: warning: literal string will be frozen in the future (run with --debug-frozen-string-literal for more information)
"hello world"

Each literal object warns once, on its first mutation, whatever the method (<<, upcase!, []=). %q() literals and heredocs are literals too and warn the same way. Strings that were never bare literals never warn: an interpolated string such as "a#{1}b", +"hello", "hello".dup, and String.new("hello") are all plain mutable strings, while -"hello" is frozen and stays so. --disable=frozen-string-literal silences the warning; --enable=frozen-string-literal or the # frozen_string_literal: true magic comment turns it into the FrozenError described in the following entry. String#+@ is the idiom for a mutable copy: on 3.4 it duplicates exactly when mutating the string would warn.

warning: literal string will be frozen in the future

text
chilled.rb:2: warning: literal string will be frozen in the future (run with --debug-frozen-string-literal for more information)

With --debug-frozen-string-literal added, the parenthetical goes away and a second line points at the literal:

text
$ ruby -W:deprecated --debug-frozen-string-literal chilled.rb
chilled.rb:2: warning: literal string will be frozen in the future
chilled.rb:1: info: the string was created here
"hello world"

Cause: The code on the reported line mutated a string that came straight from a literal in the source. Fix: Build a mutable copy where the literal is created, with +"hello", "hello".dup, or String.new("hello"), or stop mutating it and build the result another way: +, interpolation, or Array#join.

can't modify frozen String: "hello" (FrozenError)

The message quotes the string’s own content, so yours names your literal rather than "hello". Under # frozen_string_literal: true, or ruby --enable=frozen-string-literal, the chilled-string warning becomes this error:

Ruby
# frozen_string_literal: true
s = "hello"
s << " world"
text
$ ruby frozen.rb
frozen.rb:3:in '<main>': can't modify frozen String: "hello" (FrozenError)

Cause: A mutating method (<<, gsub!, []=, upcase!) was called on a frozen literal. Fix: Build a mutable copy first, with +"hello", "hello".dup, or String.new("hello"), or call the non-mutating method (+, gsub, upcase) and assign its result.

HTML, URL, and Shell Escaping Are Different Jobs

“Escaping” also names three jobs that have nothing to do with Ruby’s string syntax: making text safe for an HTML page, for a URL, and for a shell command. Each has its own method, and a backslash is the wrong tool for all three:

Ruby
require "erb"
require "cgi"
require "uri"
require "shellwords"
 
ERB::Util.html_escape(%q(<a href='x'>Tom & Jerry "Co"</a>))
# => "&lt;a href=&#39;x&#39;&gt;Tom &amp; Jerry &quot;Co&quot;&lt;/a&gt;"
CGI.escape("a b&c=d/é")                   # => "a+b%26c%3Dd%2F%C3%A9"
URI.encode_uri_component("a b&c=d/é")     # => "a%20b%26c%3Dd%2F%C3%A9"
Shellwords.escape("it's a file name.txt") # => "it\\'s\\ a\\ file\\ name.txt"
Shellwords.join(["ls", "-l", "my dir"])   # => "ls -l my\\ dir"

CGI.escapeHTML returns exactly what ERB::Util.html_escape returns. For URLs, CGI.escape and URI.encode_www_form_component encode a space as +, the form-body convention; ERB::Util.url_encode and URI.encode_uri_component encode it as %20, which is what a path segment needs. For shell commands, the best escape is none: pass arguments separately, system("echo", "it's a file; rm -rf /"), and no shell ever sees them. On Ruby 3.4 these four libraries ship as default gems (cgi 0.4.2, erb 4.0.4.1, shellwords 0.2.2, uri 1.0.4), so the require lines need no Gemfile entry.

Escaping Errors: Causes and Fixes

Every message here is copied from a Ruby 3.4.10 run. Syntax errors are shown as ruby -e one-liners: the caret diagnostics come from Prism, Ruby 3.4’s default parser, and a file run wraps the same diagnostic in extra lines. Runtime errors are shown as file runs, where error_highlight adds its caret block.

unterminated string meets end of file

text
$ ruby -e 'puts "hello'
-e: -e:1: syntax error found (SyntaxError)
> 1 | puts "hello
    |            ^ unterminated string meets end of file

The most common trigger is an apostrophe inside single quotes, which closes the string early and leaves the real closing quote to open a new one. Prism reports both problems on one line:

text
$ ruby -e "puts 'don't'"
-e: -e:1: syntax errors found (SyntaxError)
> 1 | puts 'don't'
    |           ^ unexpected local variable or method, expecting end-of-input
    |            ^ unterminated string meets end of file

Cause: The quote that opened the string never closed. Fix: Escape the apostrophe as \', or switch the delimiter: double quotes, %q(), and heredocs all hold an apostrophe as it is. A \c or \M- written right before the closing quote gives the same message, because the escape consumes the quote as its argument.

unexpected local variable or method, expecting end-of-input

text
$ ruby -e 'puts "say "hi""'
-e: -e:1: syntax error found (SyntaxError)
> 1 | puts "say "hi""
    |            ^~ unexpected local variable or method, expecting end-of-input

Cause: An unescaped double quote closed the string after say , and Ruby parsed hi as code. Fix: Put \" before each inner quote, or a different delimiter: 'say "hi"' or %(say "hi").

invalid Unicode escape sequence and too short escape sequence: \u

text
$ ruby -e 'puts "\u{110000}"'
-e: -e:1: syntax error found (SyntaxError)
> 1 | puts "\u{110000}"
    |          ^~~~~~ invalid Unicode escape sequence
text
$ ruby -e 'puts "\uZZZZ"'
-e: -e:1: syntax error found (SyntaxError)
> 1 | puts "\uZZZZ"
    |       ^~ too short escape sequence: \u
text
$ ruby -e 'puts "\xZZ"'
-e: -e:1: syntax error found (SyntaxError)
> 1 | puts "\xZZ"
    |       ^~ invalid hex escape sequence
text
$ ruby -e 'puts "\M-\M-a"'
-e: -e:1: syntax error found (SyntaxError)
> 1 | puts "\M-\M-a"
    |       ^~~~~ invalid meta escape sequence; meta cannot be repeated

Cause: \u needs exactly four hex digits, or braces holding one or more code points no higher than U+10FFFF; \x needs hex digits; \M- and \C- cannot be applied twice ("\C-\C-a" reports invalid control escape sequence; control cannot be repeated). The classic trigger is a Windows path in double quotes: "C:\Users" survives, because \U is an unknown escape and the backslash is dropped, but "C:\xUsers" does not. Fix: write the sequence correctly, double the backslash ("C:\\xUsers"), or use single quotes. Empty braces, "\u{}", are legal and produce an empty string.

unterminated heredoc; can't find string "EOS" anywhere before EOF

A heredoc needs a file to reproduce, so this one is a file run of heredoc.rb:

Ruby
puts <<EOS
  hello
  EOS
text
$ ruby heredoc.rb
heredoc.rb: --> heredoc.rb
 
unterminated heredoc; can't find string "EOS" anywhere before EOF
 
> 1  puts <<EOS
 
heredoc.rb:1: syntax error found (SyntaxError)
> 1 | puts <<EOS
    |        ^~~ unterminated heredoc; can't find string "EOS" anywhere before EOF
  2 |   hello
  3 |   EOS

Cause: The terminator is indented under a plain <<EOS, which requires it at the start of the line, or it is misspelled. Fix: <<~EOS or <<-EOS, which both allow an indented terminator, or move the terminator to column zero.

end pattern with unmatched parenthesis (RegexpError)

In a literal, Ruby catches an unbalanced group at parse time:

text
$ ruby -e 'p /a(b/'
-e: -e:1: syntax error found (SyntaxError)
> 1 | p /a(b/
    |      ^ end pattern with unmatched parenthesis

The debugging moment is the runtime version, when user input reaches a regex through interpolation or Regexp.new. Save this as interp.rb:

Ruby
input = "("
/#{input}/
text
$ ruby interp.rb
interp.rb:2:in '<main>': end pattern with unmatched parenthesis: /(/ (RegexpError)

premature end of char-class is the same error for an unmatched [, and target of repeat operator is not specified for a leading *. Cause: Text containing regex metacharacters was interpolated as it is. Fix: Regexp.escape(input) before interpolating, or Regexp.union, which escapes for you.

invalid byte sequence in UTF-8 (ArgumentError)

One line in gsub.rb:

Ruby
"\xff".gsub(/x/, "")
text
$ ruby gsub.rb
gsub.rb:1:in 'String#gsub': invalid byte sequence in UTF-8 (ArgumentError)
 
"\xff".gsub(/x/, "")
            ^^^^^^^
	from gsub.rb:1:in '<main>'

Cause: The string carries bytes that are not valid UTF-8, and a character-aware method (gsub, split, =~, encode) had to read them. A \x or \M- escape can produce such bytes ("\M-a" is a single byte, 225), and so can data read from a Latin-1 file. Fix: If the bytes are junk, "caf\xE9".scrub replaces them with �, or scrub("?") with a character you choose; if they are text in another encoding, tell Ruby which one: "caf\xE9".encode("UTF-8", "ISO-8859-1") gives "café". encode("UTF-8", invalid: :replace, undef: :replace) is the tolerant form for input you cannot trust.

incompatible character encodings: UTF-8 and BINARY (ASCII-8BIT) (Encoding::CompatibilityError)

One line in concat.rb:

Ruby
p "é" + "\xff".b
text
$ ruby concat.rb
concat.rb:1:in 'String#+': incompatible character encodings: UTF-8 and BINARY (ASCII-8BIT) (Encoding::CompatibilityError)
	from concat.rb:1:in '<main>'

Ruby 3.4 names the binary encoding BINARY (ASCII-8BIT), and the order of the two names follows the operands. Cause: One operand is UTF-8 with non-ASCII characters and the other is binary with high bytes, so no single encoding fits the result. Pure ASCII is compatible with anything, which is why "\xff".b + "abc" works. Fix: Bring both sides to one encoding before concatenating: .b on both for a byte string, or .force_encoding("UTF-8").scrub on the binary side for text. Two siblings come from encode: "é".encode("US-ASCII") raises U+00E9 from UTF-8 to US-ASCII (Encoding::UndefinedConversionError), fixed with undef: :replace or fallback: { "é" => "e" }, and "\xff".encode("UTF-16") raises "\xFF" on UTF-8 (Encoding::InvalidByteSequenceError), fixed with invalid: :replace or scrub.

Conclusion

The single-versus-double rule decides almost everything: double quotes interpret escapes and interpolation, single quotes keep the backslash. Delimiters are how you escape less: %q(), %Q(), heredocs, and %r{} each remove a class of backslashes. On Ruby 3.4, mutate a copy of a literal rather than the literal, and the chilled-string warning never appears.

Frequently asked questions

How do I escape a double quote inside a Ruby string?
Put a backslash before it: "Hello \"world\"" produces Hello "world". Or change the delimiter so nothing needs escaping: a single-quoted string, a %Q() or %() literal, or a heredoc all hold double quotes as they are. Only the quote character that opened the string ever needs a backslash.
What is the difference between single and double quoted strings in Ruby?
Double-quoted strings process escape sequences such as \n and \t and interpolate #{} expressions. Single-quoted strings keep backslashes literally, with two exceptions: \\ for a backslash and \' for a single quote. They never interpolate. Both produce ordinary String objects; the difference is only in how Ruby reads the literal.
How do I print a literal #{} without interpolation in Ruby?
Escape the hash sign: "\#{name}" prints #{name} unchanged. A single-quoted string does the same with no escape at all, because single quotes never interpolate. The same trick covers the shorthand forms, \#$global and \#@ivar. When you inspect such a string, Ruby displays a backslash it added itself; that is normal.
How do I escape special characters in a Ruby regular expression?
Inside a regex literal, put a backslash before the metacharacter: /\[world\]/ matches literal brackets, and %r{} lets you write slashes without escaping them. For text that arrives at runtime, call Regexp.escape on it before interpolating; Regexp.union escapes its string arguments automatically.
What does warning: literal string will be frozen in the future mean in Ruby 3.4?
You mutated a bare string literal. Ruby 3.4 treats such literals as chilled: still mutable, but flagged because a future version will freeze them. The warning appears only when deprecation warnings are on, via -w or -W:deprecated. Build a mutable copy with +"text", dup, or String.new, or avoid the mutation.

Published , Updated

Wondering what you can do next?

  • Share this article on social media
Tom de Bruijn

Tom de Bruijn

Tom is a developer at AppSignal, organizer, and writer from Amsterdam, The Netherlands.

All articles by Tom de Bruijn

Become our next author!

Find out more
$appsignal install

AppSignal monitors your apps

AppSignal provides insights for Ruby, Rails, Elixir, Phoenix, Node.js, Express and many other frameworks and libraries. We are located in beautiful Amsterdam. We love stroopwafels. If you do too, let us know. We might send you some!

Discover AppSignal