My learning journey into how Go handles strings under the hood.
When I started learning Go, I thought strings were one of the simplest data types. After all, a string is just a sequence of characters, right?
Well, not exactly!
As I explored how Go handles strings, I discovered some interesting concepts that changed the way I think about text in programming. Things like UTF-8 encoding, bytes versus runes, string immutability, and memory representation are not just theoretical details. They directly affect how we write correct and efficient Go programs.

In this post, I want to share my understanding of these concepts, with practical examples that helped me connect the dots.
1. What exactly is a string in Go?
In Go, a string is an immutable sequence of bytes.
That definition might sound a little strange at first. If a string represents text, why is it defined as a sequence of bytes rather than characters?
The reason is Unicode and the way modern programming languages represent text.
In the early days of computing, ASCII was widely used to represent English characters. ASCII uses values in the range of 0 to 127, which fit into a single byte.
However, the world doesn’t communicate only in English. We have accented characters, Chinese, Arabic, Hindi, emojis, and countless other symbols.
To represent text from different languages, Unicode provides a universal system of code points. UTF-8 is one of the most widely used encoding formats for representing those code points as bytes.
Go uses UTF-8 to encode text in its source code and commonly uses UTF-8-encoded strings to represent textual data.
This leads to an important distinction:
- A string is physically stored as bytes.
- Those bytes can logically represent Unicode text.
Understanding this distinction is the foundation for working with strings in Go.
2. Bytes vs. runes: What’s the difference?
Go provides two important types for working with text:
TypeDescriptionbyteAn alias for uint8, representing an 8-bit value.runeAn alias for int32, representing a Unicode code point.
A rune is what we commonly think of as a Unicode character at the code-point level.
For example, the character é has the Unicode code point U+00E9, which is decimal 233.
Let’s look at a simple example:
package main
import "fmt"
func main() {
s := "élite"
fmt.Println("String:", s)
fmt.Println("Bytes:", len(s))
fmt.Println("Runes:", len([]rune(s)))
}
Output:
String: élite
Bytes: 6
Runes: 5
Wait, why is the byte length 6 when the string contains only 5 characters?
This is where UTF-8 becomes important.
3. Understanding UTF-8 encoding
UTF-8 is a variable-length encoding format. It uses between one and four bytes to represent a Unicode code point.
For example:
| Character | UTF-8 byte length |
|-----------|-------------------|
| `A` | `1 byte` |
| `é` | `2 bytes` |
| `€` | `3 bytes` |
| `😀` | `4 bytes` |
Let’s return to our string:
s := "élite"
The string contains five Unicode code points:
- é
- l
- i
- t
- e
The four ASCII characters each require one byte, while é requires two bytes in UTF-8.
So the total is:
2 + 1 + 1 + 1 + 1 = 6 bytes.
We can verify this in Go:
package main
import "fmt"
func main() {
s := "élite"
fmt.Println(len(s)) // 6
fmt.Println(len([]byte(s))) // 6
fmt.Println(len([]rune(s))) // 5
}
Here’s what each expression does:
- len(s) returns the number of bytes in the string.
- len([]byte(s)) converts the string to a byte slice and returns its length.
- len([]rune(s)) converts the string into Unicode code points and returns the number of runes.
This is an important lesson: len() does not count Unicode characters. It counts bytes.
Whenever we work with international text, we need to keep this in mind.
4. String indexing in Go
One of the most surprising things I learned is that indexing a string returns a byte, not a rune.
Consider this example:
package main
import "fmt"
func main() {
s := "é"
fmt.Println(len(s)) // 2
fmt.Println(s[0]) // 195
fmt.Println(s[1]) // 169
}
Why does this happen?
The Unicode code point for é is decimal 233, but its UTF-8 representation consists of two bytes:
C3 A9
Those hexadecimal byte values correspond to decimal 195 and 169.
When we access s[0], Go returns the first byte of the UTF-8 encoding, not the complete Unicode code point.
If we want to work with the Unicode code point, we can convert the string to a rune slice:
s := "é"
runes := []rune(s)
fmt.Println(runes[0]) // 233
This conversion decodes the UTF-8 bytes into Unicode code points.
Iterating over strings correctly
Go provides a convenient way to iterate over the Unicode code points in a string using range.
package main
import "fmt"
func main() {
s := "élite"
for i, r := range s {
fmt.Printf(
"Byte position: %d, Rune: %c, Code point: %U\n",
i, r, r,
)
}
}
The important thing to remember is that i represents the byte offset, while r is the rune value.
For multi-byte characters, the byte offset increases by more than one between iterations.
This is generally the approach I would use when I need to process Unicode code points rather than individual bytes.
One additional detail: a rune is not always the same as a character perceived by a user. Some visible characters, such as certain emojis or accented letters, can consist of multiple Unicode code points. For those cases, we need grapheme-cluster-aware processing.
5. How Go Strings Are Represented in Memory
To understand Go strings, imagine a string as a window into a sequence of bytes.
Conceptually, a Go string contains two things:
- A pointer to the beginning of the underlying data.
- A length representing the number of bytes.
For example:
s := "hello, world"
fmt.Println(len(s)) // 12
Since every character here is ASCII, the string contains 12 bytes.
Unlike C-style strings, Go strings don’t require a null terminator (\0). Go already knows the string's length, so len(s) is an O(1) operation.
5.1 Substrings: A Window Into Existing Data
Consider this example:
s := "hello, world"
hello := s[:5]
world := s[7:]
fmt.Println(hello) // hello
fmt.Println(world) // world
You might assume Go creates new copies of the bytes for hello and world. But in Go's standard implementation, string slicing can share the original underlying byte storage.

String slicing creates new string views that can share the same underlying bytes.
As the diagram shows, the original string has a data pointer and a length of 12. The substring hello starts at byte 0 and has length 5, while world starts at byte 7 and has length 5.
The bytes don’t need to move or be copied.
Mental model: A substring is like a window over existing data. It changes where the view starts and how much data it covers, not necessarily the underlying storage.
5.2 Why Can a Small Substring Keep a Large String in Memory?
This sharing can have an interesting consequence.
Imagine the original string contains several megabytes of data, but you only need five characters:
small := large[:5]
large = ""
Even after large is no longer referenced, small may still point into the original allocation. The garbage collector may therefore keep the entire allocation reachable.
A five-byte substring can effectively keep a much larger block of memory alive.
5.3 Using strings.Clone
If you want an independent copy of the substring, use strings.Clone:
package main
import (
"fmt"
"strings"
)
func main() {
large := "hello, world"
small := strings.Clone(large[:5])
fmt.Println(small) // hello
}
strings.Clone copies the substring's bytes into independent storage.
| Without `strings.Clone` | With `strings.Clone` |
|-------------------------------------------|---------------------------------------------------|
| Substring may share the original storage. | Substring has its own copy. |
| May keep a larger allocation reachable. | Original storage can be reclaimed if otherwise unused. |
| Avoids copying the bytes. | Requires copying and may allocate memory. |
Use strings.Clone when you need an independent copy, particularly when retaining a small substring from a large string for a long time.
So a Go string’s logical length is not necessarily the size of the memory allocation it keeps reachable.
Understanding this distinction explains why string slicing is efficient and why strings.Clone can sometimes help reduce unnecessary memory retention.
6. Strings are immutable in Go
Another important characteristic of Go strings is immutability.
Once a string value has been created, we cannot modify its individual bytes through string indexing.
For example, this is not allowed:
s := "hello"
s[0] = 'H' // Compile-time error
We cannot change the first character of an existing string this way.
Instead, we create a new string value and assign it to the variable.
s := "hello"
s = "Hello"
fmt.Println(s) // Hello
The variable s now holds the new string value.
What happens when we assign a string?
Consider:
s := "quick brown fox"
t := s
s += "es"
fmt.Println(s) // quick brown foxes
fmt.Println(t) // quick brown fox
When we assign s to t, the string value is copied. The underlying bytes can be shared because strings are immutable.
When we concatenate "es" to s, the result is assigned back to s. The value held by t remains unchanged.
This demonstrates an important distinction:
Changing a string variable is not the same as modifying the contents of an existing string.
The same principle applies to string functions.
For example:
package main
import (
"fmt"
"strings"
)
func main() {
s := "hello"
upper := strings.ToUpper(s)
fmt.Println(s) // hello
fmt.Println(upper) // HELLO
}
The strings.ToUpper function returns a string containing the uppercase result. It does not modify the original string.
If we want to update the variable, we can assign the result back:
s = strings.ToUpper(s)
Because strings are immutable, the Go runtime can safely share string data between values. When data is no longer reachable, it can eventually be reclaimed by the garbage collector.
The exact allocation and copying behavior depends on the implementation, but immutability itself is guaranteed by the language.
7. Working with strings: A simple search-and-replace program
Let’s put some of these concepts into practice by building a small command-line program that replaces occurrences of one string with another.
The program will:
- Accept the old and new strings as command-line arguments.
- Read input from standard input, one line at a time.
- Replace occurrences of the old string.
- Print the modified lines.
Here’s the implementation:
package main
import (
"bufio"
"fmt"
"os"
"strings"
)
func main() {
if len(os.Args) < 3 {
fmt.Println("Usage: go run main.go <old> <new>")
return
}
old := os.Args[1]
new := os.Args[2]
scanner := bufio.NewScanner(os.Stdin)
for scanner.Scan() {
line := scanner.Text()
parts := strings.Split(line, old)
result := strings.Join(parts, new)
fmt.Println(result)
}
if err := scanner.Err(); err != nil {
fmt.Fprintln(os.Stderr, "Error:", err)
}
}
Suppose we have an input file called test.txt:
Matt went to Greece.
Matt likes Go.
Matt is learning strings.
We can run the program like this:
go run main.go Matt Ed < test.txt
Output:
Ed went to Greece.
Ed likes Go.
Ed is learning strings.
How does Split and Join work?
Let’s understand the core logic with a smaller example.
line := "Matt went to Matt's class"
parts := strings.Split(line, "Matt")
fmt.Println(parts)
The result is:
[ went to 's class]
More precisely, the slice contains three strings:
[]string{"", " went to ", "'s class"}
The original separator, "Matt", is removed during the split.
Now let’s join the parts using "Ed":
result := strings.Join(parts, "Ed")
fmt.Println(result)
Output:
Ed went to Ed's class
This works because strings.Join inserts the specified separator between each pair of adjacent elements in the slice.
For simple replacement tasks, Go also provides a more direct function:
result := strings.ReplaceAll(line, "Matt", "Ed")
This is generally easier to read when the only goal is replacing every occurrence of a substring.
The example is still useful for understanding how strings, slices, standard input, and command-line arguments work together.
A small practical note: bufio.Scanner has a default maximum token size. If we're processing very long lines, we may need to increase its buffer size or choose another input-reading approach.
8. A few useful string operations in Go
Go’s standard strings package provides many useful functions for everyday programming.
Here are a few that I expect to use frequently:
| Function | Purpose |
|--------------------------------------|-----------------------------------------------------------|
| `strings.Contains(s, substr)` | Checks whether a string contains another string. |
| `strings.HasPrefix(s, prefix)` | Checks whether a string starts with a given prefix. |
| `strings.HasSuffix(s, suffix)` | Checks whether a string ends with a given suffix. |
| `strings.ToUpper(s)` | Returns an uppercase version of a string. |
| `strings.ToLower(s)` | Returns a lowercase version of a string. |
| `strings.Split(s, sep)` | Splits a string around a separator. |
| `strings.Join(slice, sep)` | Joins a slice of strings using a separator. |
| `strings.ReplaceAll(s, old, new)` | Replaces all non-overlapping occurrences of a substring. |
| `strings.TrimSpace(s)` | Removes leading and trailing whitespace. |
These functions make common text-processing tasks much easier without having to manually manipulate bytes.
For Unicode-aware processing, it’s still important to understand what each function considers a character, a byte, or a substring.
9. Key takeaways from my learning
After exploring Go strings, here are the concepts I’m taking away:
1. Strings are sequences of bytes. Go strings commonly contain UTF-8-encoded text, but they are not restricted to valid UTF-8.
2. Bytes and runes are different. A byte is an 8-bit value, while a rune represents a Unicode code point.
3. len(s) returns bytes, not characters. When working with Unicode code points, len([]rune(s)) provides a different count, though it still doesn't necessarily count user-perceived characters.
4. String indexing operates on bytes. Use range or convert to []rune when you need to process Unicode code points.
5. Strings are immutable. String operations produce values rather than modifying existing string contents in place.
6. String slicing can share underlying storage. This can improve efficiency, but it is useful to understand the implications when working with large strings.
7. Go provides powerful string utilities. The strings package makes searching, splitting, joining, and replacing text straightforward.
Conclusion
Strings may look simple on the surface, but understanding how they work internally makes a big difference when writing Go programs.
The most important mental model I’m taking away is that a Go string has two sides: its physical representation as bytes and its logical interpretation as Unicode text.
Once we understand the difference between bytes, runes, and user-perceived characters, many otherwise confusing behaviors become much easier to explain.
I’m continuing to explore Go one concept at a time, and this was a useful reminder that understanding the fundamentals often makes the more advanced topics much easier.
What about you? Have you ever run into unexpected behavior when working with Unicode strings in Go? I’d love to hear about your experiences and the approaches you use to handle text correctly.
#Go #Golang #Programming #LearnInPublic #SoftwareEngineering