EMZETT.
Login

String

In short: A data type for character sequences (text) in programming.

In more detail: Consists of a sequence of characters (letters, numbers, symbols), usually stored internally as an array of characters or bytes — the specific character encoding is governed by, for example, Unicode. In many languages, strings are immutable: every change produces a new string instead of modifying the existing one.

In Depth

Immutability is an important, often surprising detail: if you write, say, text = text + "!" in Java, the existing string in memory is NOT modified — a completely new string is created and the variable then points to this new one — the old string remains unchanged in memory until removed by the garbage collector. This makes strings safe in concurrent code (several threads can read the same string without interfering with each other), but costs performance for very many concatenations in a loop — for that, there are specifically mutable alternatives like StringBuilder in Java.

Typical string operations are similar across almost all languages: getting the length, extracting substrings, searching/replacing, changing case, splitting at a delimiter (split), and joining several parts together (join/concatenation). How a single character is internally encoded as a number is governed by Unicode — modern languages treat strings as Unicode text by default, correctly supporting umlauts, emoji and non-Latin writing systems, unlike older, ASCII-limited systems.

Comparing strings: reference vs. content

A common beginner mistake concerns comparing strings: in many languages, == on objects (which strings internally count as) by default checks whether two variables point to the same object in memory (“reference equality”), not whether their content is equal. Java, for example, explicitly distinguishes between == (reference comparison) and .equals() (content comparison) — two strings with identical content but created differently in memory can appear “unequal” with ==, even though they contain the same text. Other languages like Python or JavaScript solve this differently (JavaScript’s === compares primitive strings by content, not reference), which often causes confusion when switching languages.

Escape sequences

To represent special characters within a string that would otherwise collide with the syntax (e.g. a quotation mark inside a string enclosed in quotation marks), practically all languages use escape sequences with a backslash: \n for a line break, \t for a tab, \" for a literal quotation mark, \\ for a literal backslash itself. This convention is largely identical across almost all C-influenced languages (Java, JavaScript, Python, C itself), which makes switching between languages easier on this point.

Template/interpolated strings

Besides classic concatenation ("Hello " + name), modern languages usually also offer a more compact interpolation syntax, to embed variables directly into a string: template literals in JavaScript (`Hello ${name}`), f-strings in Python (f"Hello {name}"), string templates in Kotlin ("Hello $name"). This syntax is purely syntactic sugar — internally, a concatenation/formatting is still performed — but considerably improves readability compared to long chains of + concatenations, especially when several variables are embedded in a text.

See also: Unicode, Strings