Commit Graph

7 Commits

Author SHA1 Message Date
gffranco c8a6fc1172 parity(8b): multi-line list item continuation
Lines indented strictly past a list item's marker now attach to that
item rather than becoming a sibling paragraph. Matches upstream
vimwiki + Markdown convention.

Before:
  - first item
    continuing here          ←  rendered as a separate <p>
  - second

Now:
  - first item continuing here  ← one <li>, joined with a space
  - second

The detection lives in `parse_list_item`: after the item's first
line, we loop over subsequent lines and treat each as continuation
when its leading whitespace exceeds the marker's column. Two
flavours need handling because the lexer emits different tokens
based on indent width:

  * 1..3 leading spaces → `Text(s)` with the whitespace embedded;
    we trim it off the first text token before re-running
    `parse_inline_seq`.
  * 4+ leading spaces  → `BlockquoteIndent` + bare `Text`; we
    advance past the indent token and use the text as-is.

Continuation also works inside nested lists (`level + 1` threshold
adapts to the marker's own indent) and is bounded by blank lines /
unindented content / sibling-or-shallower list markers.

Tests: 6 new in `crates/nuwiki-core/tests/list_continuation.rs`
covering single-line continuation, multi-line continuation, blank-
line termination, unindented termination, nested-list continuation,
and the HTML-renderer regression (single `<ul>` with two `<li>`s,
no orphan `<p>`).

Gates: 456 Rust / 39 Neovim / 12 Vim.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 22:15:58 +00:00
gffranco bcef714805 parity(8a): transclusion attr quote-stripping
`{{url|alt|key="quoted value"}}` syntax was already parsed end-to-end
through to `TransclusionNode::attrs`, but the value retained the
surrounding `"…"` (or `'…'`) characters verbatim. The HTML renderer
then escaped them into `&quot;`, producing `style="&quot;border: 1px&quot;"`
instead of clean `style="border: 1px"`.

Extracted a small `insert_attr` helper that splits on `=`, trims, and
strips matching single- or double-quote pairs from the value before
inserting. Unquoted values pass through unchanged.

Tests: 4 new in crates/nuwiki-core/tests/transclusion_attrs.rs
covering double-quote stripping, single-quote stripping, unquoted
pass-through, and the clean-HTML rendering shape (no `&quot;`).

Note: this commit covers the "transclusion attrs" half of the
original Cluster 8 audit. The "multi-line list-item continuation
paragraphs" half requires lexer + parser changes (continuation lines
indented past the marker should attach to the prior list item
instead of becoming a sibling paragraph) and is deferred.

Gates: 429 Rust / 39 Neovim / 12 Vim.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 21:51:54 +00:00
gffranco 214d54e3b9 parity(7): table cell alignment markers (|:--|--:|:--:|)
The Markdown-style alignment anchors on a table's header-separator
row now propagate from the lexer through the parser into the AST,
and the HTML renderer emits a `style="text-align: …;"` attribute on
each affected `<th>` / `<td>`.

  |:--|   → left
  |--:|   → right
  |:--:|  → center
  |---|   → default (no style attr)

Encoding choice: rather than threading the source string into the
parser, the lexer now extracts per-cell alignment when it recognises
the separator row and carries the `Vec<TableAlign>` inline on the
`TableHeaderRow` token. The parser unpacks it into `TableNode`'s
new `alignments: Vec<TableAlign>` field. Cells past the end of the
vector render with the default alignment.

`TableAlign` lives in `nuwiki_core::ast::block` alongside the table
nodes and is re-exported via `nuwiki_core::ast`.

Tests: 4 new in crates/nuwiki-core/tests/vimwiki_table_alignment.rs
covering plain dashes, all three anchor flavours, the HTML output
shape, and the no-style-attr case. The existing
`vimwiki_lexer::table_header_separator_row` was updated to assert
the new payload shape.

Gates: 425 Rust / 39 Neovim / 12 Vim.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 21:45:43 +00:00
gffranco 8b29d4e9b9 phase 12: tags — lexer, AST, index, anchor resolution, ws/symbol
CI / cargo fmt --check (push) Successful in 18s
CI / cargo clippy (push) Successful in 1m3s
CI / cargo test (push) Successful in 1m11s
Vimwiki `:tag1:tag2:` syntax across the full stack. P11 resolved as
strict-vimwiki — no markdown `#tag` alias.

AST (nuwiki-core):
- `TagScope { File, Heading(usize), Standalone }` mirrors vimwiki's
  three placement-determined meanings: lines 0-1 → File; one or two
  lines below the most recent heading → Heading(idx); otherwise
  Standalone.
- `TagNode { span, tags: Vec<String>, scope }` + `BlockNode::Tag`.
- `PageMetadata.tags: Vec<String>` accumulates the file-scope tag
  names — convenience projection of the Tag blocks whose scope is
  File. New field is additive (Default derive) so existing consumers
  using `..Default::default()` keep working.
- `Visitor::visit_tag` lands with a default body and `walk_block`
  dispatches the new arm.

Lexer (`syntax/vimwiki/lexer.rs`):
- `VimwikiTokenKind::Tag(Vec<String>)` block-level token.
- `try_lex_tag` runs in the normal-line dispatcher right after heading
  recognition; demands the *whole trimmed line* match `:name:name…:`,
  with non-empty whitespace-free names between adjacent colons. Empty
  segments (`::tag::`), whitespace inside a name, or single-colon
  lines are all rejected. Doesn't fight `Term::` (definition-list
  marker) because that pattern requires text before the `::`.
- Indentation is tolerated; the span covers the trimmed run.

Parser (`syntax/vimwiki/parser.rs`):
- `ParseState` gains `headings_seen` + `last_heading_line`.
- `parse_heading` bumps both on every closed heading so the tag
  parser can detect the heading window.
- `parse_tag` reads the Tag token, computes the scope via
  `scope_for_tag(tag_line)`, eats the trailing newline, and emits the
  TagNode.
- `parse_document` post-processes each block: when a Tag with
  `scope == File` is emitted, its names are pushed (de-duped) into
  `metadata.tags`.

Renderer (`render/html.rs`):
- `render_tag` emits `<div class="tags"><span class="tag"
  id="tag-name">name</span>...</div>`. The `id` is what makes
  `[[Page#tag-name]]` jump to the rendered tag.

Semantic tokens (`semantic_tokens.rs`):
- New `TOKEN_TAG = 19` / `vimwikiTag` so editors highlight tag lines
  distinctly. `emit_block` adds the Tag arm.

Workspace index (`index.rs`):
- `IndexedPage.tags: Vec<TagInfo>` records each tag occurrence on a
  page with its span + scope.
- `WorkspaceIndex.tags_by_name: HashMap<String, Vec<TagOccurrence>>`
  is the reverse map for cross-page anchor resolution + the Phase 13
  `nuwiki.tags.search` command.
- `upsert` populates both; `remove` scrubs the reverse map and drops
  empty entries; `tags_for(name)` is the public lookup.

LSP handlers (`lib.rs`):
- `heading_range_in` falls back to tag matching when the anchor
  isn't a heading slug — so `[[Page#some-tag]]` lands on the matching
  TagNode. Heading anchors win over tag anchors when both match (the
  vimwiki precedence).
- `workspace/symbol` includes tags alongside headings, named `:tag:`
  and kinded `PROPERTY` so editor UIs can distinguish them.

Tests (20 new):
- Syntax (15 in `vimwiki_tags.rs`):
  lexer recognition + rejection of `::tag::`, whitespace-in-name,
  Term::-conflict; parser scope paths (file, heading-1-below,
  heading-2-below, standalone, no-heading, line-1 priority);
  metadata dedup; renderer HTML + escaping.
- LSP index (5 in `tags_index_and_lsp.rs`):
  per-page tag population, cross-page `tags_by_name` aggregation,
  re-upsert replaces stale tags, `remove` cleans the reverse map,
  heading-scope tag indexed but not in file metadata.

All 191 v1.0 + Phase 11 tests still pass — backward compatible.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 14:54:55 +00:00
gffranco 294963b517 phase 4: vimwiki parser + SyntaxPlugin glue
CI / cargo fmt --check (push) Successful in 19s
CI / cargo clippy (push) Successful in 1m2s
CI / cargo test (push) Successful in 1m5s
Hand-rolled recursive-descent parser over the token stream. Resilient
per SPEC §6.7: never aborts the document — anything the dispatcher
can't claim becomes BlockNode::Error and the cursor advances, and
unterminated wikilinks / transclusions / delimiters fall back to
literal text.

Block coverage: every variant from §6.4 — Heading (with inline
content between markers), Paragraph (with soft-break across single
newlines), HorizontalRule, Blockquote (both `>` and 4-space styles),
Preformatted, MathBlock, List (with checkbox + indent-driven
sublists), DefinitionList, Table (header separator promotes the
preceding row, plus colspan / rowspan cells), Comment (single +
multiline), and page-level placeholders threaded into PageMetadata.

Inline pairing: left-to-right scan with "find next matching delim",
recursing on the slice between. `_*x*_` becomes Italic(Bold(...)),
mirroring SPEC §6.7 precedence. URLs inside `[[ ]]` route to
ExternalLinkNode; everything else to WikiLinkNode with full
LinkTarget classification (Wiki / Interwiki numbered + named /
Diary / File / Local / AnchorOnly / directory / root-relative /
filesystem-absolute).

VimwikiSyntax registers as id="vimwiki", extensions=[".wiki"] and
chains lexer + parser inside SyntaxPlugin::parse.

SPEC §4 updated to record the hand-rolled choice with rationale —
span-bearing tokens fit awkwardly into winnow/nom combinator style;
resilience and recovery are explicit in code instead.

Tests (41 new parser cases) cover every AST construct, inline
precedence including nested bold/italic, every LinkKind, transclusion
alt + attrs, resilience on unterminated links, and end-to-end
SyntaxRegistry dispatch.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 17:58:41 +00:00
gffranco 942dbe2aa8 phase 3: vimwiki lexer
CI / cargo fmt --check (push) Successful in 22s
CI / cargo clippy (push) Successful in 53s
CI / cargo test (push) Successful in 1m2s
Two-pass, hand-rolled lexer per SPEC.md §6.6. Block pass walks the
source line by line; recognised constructs emit structural tokens, and
text-bearing line content is fed to the inline pass. Multi-line fences
(`{{{ }}}`, `{{$ }}$`, `%%+ +%%`) flip a `BlockMode` so subsequent
lines are captured raw until the matching closer.

VimwikiToken covers the full §9 feature checklist:

- headings (1–6, centered), horizontal rule, blockquotes (`>` and
  4-space), all list-marker variants, checkboxes, definition `::`,
  tables (separators, header row, col/row span)
- preformatted fences (with language + key=val attrs), math blocks
  (with `%env%`), single- and multi-line comments, all four
  page placeholders (%title / %nohtml / %template / %date)
- inline: bold/italic/strike/super/sub delimiters, inline code, inline
  math, all six keywords (word-bounded), wikilinks (with description
  separator), transclusions (with attr separator), raw URLs across
  http(s)/ftp/mailto/file schemes — trailing sentence punctuation
  stripped from URL spans

Spans are byte-accurate per SPEC §6.5 (0-indexed line + byte column +
absolute byte offset). The lexer is permissive: every delimiter is
emitted, the parser will pair them and fall back to literal text on
mismatches.

Tests (41 cases) cover one example per construct in §9, plus
word-boundary keyword detection, list-vs-bold disambiguation,
indent-vs-list disambiguation, and a span-correctness check.

README now carries a per-phase status table; SPEC top-line status
advanced to Phase 3 complete.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 17:20:46 +00:00
gffranco 8f35c0a80c phase 2: syntax plugin interface and registry
CI / cargo fmt --check (push) Successful in 19s
CI / cargo clippy (push) Successful in 52s
CI / cargo test (push) Successful in 51s
Add the Lexer/Parser/SyntaxPlugin traits and the SyntaxRegistry per
SPEC.md §6.3.

- `Lexer<Token>` and `Parser<Token>` are typed building blocks: each
  syntax owns its own token alphabet (vimwiki today, markdown later).
- `TokenStream<T>` is a `Vec<T>` newtype so the eager-vs-lazy choice
  (§6.6) can flip later without breaking callers.
- `SyntaxPlugin` is type-erased: it exposes id / display_name /
  file_extensions / parse(text) so the registry can hold heterogeneous
  plugins behind a single trait object. Implementations chain their
  Lexer + Parser inside parse().
- `SyntaxPlugin: Send + Sync` so the LSP server can stash plugins in
  an Arc and share them across handler tasks.
- `SyntaxRegistry` stores `Arc<dyn SyntaxPlugin>`, supports lookup by
  id and by file extension, iteration in registration order.

Tests cover TokenStream round-trip, full lex→parse chain through a
mock plugin, registry lookup hits and misses, and a compile-time
assert_send_sync::<SyntaxRegistry>().

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 16:58:00 +00:00