Expand description
Position-encoding conversion (LSP 3.17 §3.17 / §General).
Lattice uses UTF-8 byte offsets internally
(lattice_protocol::Position { line, byte }). LSP sends:
- utf-8 (
PositionEncodingKind::UTF8): column is the UTF-8 byte offset within the line. Identical to ours – no conversion needed. - utf-16 (
PositionEncodingKind::UTF16, the LSP 3.16 default): column is the UTF-16 code-unit offset within the line. We convert when the negotiated encoding is utf-16. - utf-32 (
PositionEncodingKind::UTF32): column is the Unicode codepoint count. Practically unused; we don’t advertise support, but the converter is here in case a server demands it.
The converters work line-by-line: callers pass the line text
plus an offset and get back the offset in the target
encoding. Lattice’s Position::byte is always within a
single line, so we never have to walk multiple lines.
§Performance
utf8_byte_to_utf16_column is O(byte) – it walks the
prefix and counts UTF-16 code units. For ASCII lines this
collapses to byte (no multi-byte chars). The bench
lsp::position::utf8_to_utf16 measures the worst case
(line of CJK glyphs) at sub-microsecond.
Functions§
- byte_
to_ lsp_ character - Convert a UTF-8 byte offset within
lineto the offset in the negotiatedencoding. Used when constructing LSPPosition::characterfrom lattice’sPosition::byte. - lsp_
character_ to_ byte - Convert an LSP
charactervalue (in the negotiated encoding) to a UTF-8 byte offset withinline. Used for ranges that arrive FROM the server (definitions, diagnostics, etc.). - utf8_
byte_ to_ utf16_ column - UTF-8 byte offset → UTF-16 code-unit offset within
line.byteis treated as a position within the line text; if it’s past the end the function returns the line’s full utf-16 length plus the over-shoot in bytes (a useful approximation for clamped callers, but production paths shouldn’t pass beyondline.len()). - utf8_
byte_ to_ utf32_ column - UTF-8 byte offset → UTF-32 codepoint offset within
line. - utf16_
column_ to_ utf8_ byte - UTF-16 code-unit offset → UTF-8 byte offset within
line. Walks chars accumulating utf-16 units; stops when the running count reachescharacter. - utf32_
column_ to_ utf8_ byte - UTF-32 codepoint offset → UTF-8 byte offset.