Skip to main content

Module position

Module position 

Source
Expand description

Position-encoding conversion (LSP 3.17 §3.17 / §General).

Lattice uses UTF-8 byte offsets internally (lattice_protocol::Position { line, byte }). LSP sends:

  • utf-8 (PositionEncodingKind::UTF8): column is the UTF-8 byte offset within the line. Identical to ours – no conversion needed.
  • utf-16 (PositionEncodingKind::UTF16, the LSP 3.16 default): column is the UTF-16 code-unit offset within the line. We convert when the negotiated encoding is utf-16.
  • utf-32 (PositionEncodingKind::UTF32): column is the Unicode codepoint count. Practically unused; we don’t advertise support, but the converter is here in case a server demands it.

The converters work line-by-line: callers pass the line text plus an offset and get back the offset in the target encoding. Lattice’s Position::byte is always within a single line, so we never have to walk multiple lines.

§Performance

utf8_byte_to_utf16_column is O(byte) – it walks the prefix and counts UTF-16 code units. For ASCII lines this collapses to byte (no multi-byte chars). The bench lsp::position::utf8_to_utf16 measures the worst case (line of CJK glyphs) at sub-microsecond.

Functions§

byte_to_lsp_character
Convert a UTF-8 byte offset within line to the offset in the negotiated encoding. Used when constructing LSP Position::character from lattice’s Position::byte.
lsp_character_to_byte
Convert an LSP character value (in the negotiated encoding) to a UTF-8 byte offset within line. Used for ranges that arrive FROM the server (definitions, diagnostics, etc.).
utf8_byte_to_utf16_column
UTF-8 byte offset → UTF-16 code-unit offset within line. byte is treated as a position within the line text; if it’s past the end the function returns the line’s full utf-16 length plus the over-shoot in bytes (a useful approximation for clamped callers, but production paths shouldn’t pass beyond line.len()).
utf8_byte_to_utf32_column
UTF-8 byte offset → UTF-32 codepoint offset within line.
utf16_column_to_utf8_byte
UTF-16 code-unit offset → UTF-8 byte offset within line. Walks chars accumulating utf-16 units; stops when the running count reaches character.
utf32_column_to_utf8_byte
UTF-32 codepoint offset → UTF-8 byte offset.