Built-in search plugin¶
MaterialX for MkDocs 10.2.0 introduces a completely redesigned search system.
Search engines are now implemented as interchangeable providers, giving the built-in search interface
better result quality, more efficient indexing, and more flexible configuration.
The new architecture is particularly well suited to large sites and adds specialized support for Chinese and Japanese content.
Pagefind is the default provider for sites served over HTTP.
Lunr remains available for sites built with the offline plugin that must also work when opened directly from file://.
Both providers are fully client-side and require no hosted search service.
Objective¶
How it works¶
The plugin builds the selected provider's index after MkDocs has rendered the site. The frontend loads that provider through a common integration and renders the results with MaterialX's built-in search interface.
Pagefind scans the generated HTML and writes a chunked index beside the site.
At search time, the browser loads only the chunks and result data needed for the current query.
Lunr writes search_index.json, then constructs and queries its in-memory index in a Web Worker.
Integration with other plugins¶
The search plugin integrates with other built-in plugins:
-
The offline plugin makes it possible to distribute the generated
sitedirectory as a.zipfile. Use it together with the Lunr provider so search also works from the local filesystem. -
The meta plugin can apply search metadata to a complete documentation section, making it easy to exclude a folder from either provider or tune Lunr result ranking.
Configuration¶
search – built-in
The search plugin is built into MaterialX for MkDocs and doesn't need to be
installed. Add search to the plugins list to enable it. Pagefind is the
default provider:
Configuration structure¶
The following example shows the provider-based structure and the options that are useful for most customizations. All settings below it are optional.
plugins:
- search:
provider: pagefind # pagefind (default) or lunr
pagefind:
# Index configuration
include_characters: ._
keep_index_url: true
exclude_selectors:
- .md-banner
# output_subdir: search
# logfile: pagefind.log
# Browser Search API configuration
options:
excerptLength: 30
exactDiacritics: false
ranking:
termFrequency: 1.0
termSimilarity: 1.0
pageLength: 0.75
termSaturation: 1.4
diacriticSimilarity: 0.8
metaWeights:
title: 5.0
lunr:
lang:
- en
separator: '[\s\-]+'
pipeline:
- stemmer
- stopWordFilter
- trimmer
# jieba_dict: dict.txt
# jieba_dict_user: user_dict.txt
Start with plugins: [search] and add only the settings required by the site.
The provider sections below explain when each option is useful.
Provider¶
Use this setting to select the search provider:
Provider-specific settings are isolated under their matching key, so Pagefind settings don't affect Lunr and vice versa.
Keep the default Pagefind provider when the site is served over HTTP or HTTPS.
When building with the offline plugin,
select Lunr if the generated site must also support search when opened directly from file://.
Pagefind¶
Pagefind is a static search library designed around a small initial payload and on-demand index loading. MaterialX installs Pagefind's extended release, which provides specialized segmentation for Chinese, Japanese, and Korean content.
Features¶
- High-quality ranking – Pagefind combines term frequency and similarity, page length, term saturation, diacritics, and content and metadata weights. This produces more relevant results across pages of different lengths and content types, while keeping the ranking model configurable.
- Precise, informative results – Results are grouped by page and matching subsection, link directly to the relevant heading, and include an excerpt around the match. Users can identify the right result before opening it.
- Chunked index – The generated index is split into small chunks instead of being delivered as one complete site-wide index.
- On-demand loading – MaterialX preloads the chunks likely to match the current query and loads full result data only when it is rendered.
- Large-site scalability – Index chunking keeps the initial payload small as content grows, making Pagefind suitable for sites with tens of thousands of pages.
- Multilingual search – Pagefind detects the language of generated pages, builds separate language indexes, and selects the appropriate index in the browser. MaterialX uses the extended release for specialized Chinese, Japanese, and Korean segmentation.
Pagefind must be served over HTTP or HTTPS because its JavaScript, WebAssembly, and index chunks are loaded on demand.
Index configuration¶
Settings directly under pagefind configure the index generated after the
MkDocs build. The following options cover the most common customizations:
| Setting | Default | Description |
|---|---|---|
exclude_selectors |
nav, footer |
CSS selectors and their descendants to omit from indexing. |
include_characters |
._ |
Punctuation preserved as searchable characters. |
keep_index_url |
true |
Keep index.html at the end of result URLs. |
logfile |
none | Also write logs to a file; relative paths are resolved inside output_subdir. |
options |
{} |
Browser Search API configuration, described in the next section. |
MaterialX marks the main content with data-pagefind-body and manages the
index input, output, and result URL format. Other Pagefind index options remain
available for advanced use; see the official index configuration
options.
Search API configuration¶
Settings under pagefind.options are passed to Pagefind's browser Search API
using camel-case names. In normal use, only excerpt, diacritic matching, and
ranking behavior need to be customized:
| Option | Default | Description |
|---|---|---|
excerptLength |
30 |
Maximum target length for generated result excerpts. |
exactDiacritics |
false |
Treat accented and unaccented characters as distinct. |
ranking |
Pagefind defaults | Tune result ranking with the parameters below. |
MaterialX manages bundle routing, result URLs, and highlighting. The complete upstream option set remains available for special cases in Pagefind's Search API configuration.
The most useful ranking controls are:
| Ranking option | Default | Purpose |
|---|---|---|
termFrequency |
1.0 |
Balance term frequency against weighted term count. |
termSimilarity |
1.0 |
Prefer indexed terms whose length is closer to the query. |
pageLength |
0.75 |
Control how strongly shorter-than-average pages are favored. |
termSaturation |
1.4 |
Control how quickly repeated terms stop increasing relevance. |
diacriticSimilarity |
0.8 |
Boost exact diacritic matches when normalization is enabled. |
metaWeights |
title: 5.0 |
Weight matches in title or custom metadata fields. |
For value ranges and the remaining controls, see Pagefind's ranking documentation.
MaterialX doesn't impose a separate schema on options.
Excluding content¶
Excluding an entire page¶
Use search.exclude in the Markdown front matter to remove a complete page and
all of its sections from the Pagefind index:
The meta plugin can apply the same property to every page in a folder:
Excluding certain types of elements¶
Use exclude_selectors for elements that should be ignored throughout the
site, such as a custom banner or generated utility block:
The matched element and all of its children are excluded.
Excluding part of a page¶
Use Pagefind's data-pagefind-ignore attribute to exclude part of a document.
For a complete section, wrap the heading and its content in an HTML container:
<div data-pagefind-ignore markdown>
## Internal notes
This complete section is excluded from Pagefind.
</div>
The attribute excludes the element and all of its children. Placing it only on a heading doesn't exclude the content that follows, which is why a wrapper is required for complete sections.
Other Pagefind features¶
Pagefind also supports metadata, filters, sorting, content weighting, custom records, and search across multiple sites. These are upstream Pagefind capabilities; some require custom templates, indexing code, or frontend code beyond MaterialX's built-in integration. See the Pagefind documentation and multisite search guide for the complete workflows.
Lunr¶
Lunr builds one JSON search index and loads it into a Web Worker in the browser. It is less efficient for very large sites than Pagefind, but remains the right provider when a generated site must work without an HTTP server.
Language¶
lang¶
Use this setting to specify the language of the search index, enabling stemming support for languages other than English. The default is computed from the site language, but can be set to one or multiple languages:
Language support is provided by lunr languages, a collection of language-specific stemmers and stop words for Lunr maintained by the Open Source community.
The following languages are currently supported by lunr languages:
ar– Arabicda– Danishde– Germandu– Dutchen– Englishes– Spanishfi– Finnishfr– Frenchhi– Hindihu– Hungarianhy– Armenianit– Italianja– Japanesekn– Kannadako– Koreanno– Norwegianpt– Portuguesero– Romanianru– Russiansa– Sanskritsv– Swedishta– Tamilte– Teluguth– Thaitr– Turkishvi– Vietnamesezh– Chinese
If lunr languages doesn't support the selected site language, the plugin falls back to the language that is expected to yield the best stemming results.
Tokenization¶
separator¶
Use this setting to specify the separator used to split words when building the search index. The default is computed from the site language, but can be set explicitly:
plugins:
- search:
provider: lunr
lunr:
separator: '[\s\-,:!=\[\]()"/]+|(?!\b)(?=[A-Z][a-z])|\.(?!\d)|&[lg]t;'
Separators support positive and negative lookahead assertions, allowing precise control over how words are split.
Broken into its parts, this separator induces the following behavior:
This inserts token boundaries before and after whitespace, hyphens, commas, brackets, and other special characters. Adjacent separators are treated as one.
This splits programming identifiers at case changes, tokenizing
PascalCase into Pascal and Case.
The negative lookahead prevents version strings like 1.2.3 from being
split into separate numbers and keeps them discoverable.
Pipeline¶
pipeline¶
Use this setting to specify the pipeline functions that filter and expand tokens after tokenization and before adding them to the index. The default is computed from the site language:
The following pipeline functions can be used:
stemmer– Stem tokens to their root form, e.g.runningtorunstopWordFilter– Filter common words such asaandthetrimmer– Trim whitespace from tokens
Segmentation¶
The plugin supports Chinese text segmentation with jieba. Japanese and Korean are segmented on the client side when Lunr is selected.
jieba_dict¶
Use this setting to specify a custom dictionary for jieba, replacing its default dictionary:
The following dictionaries are provided by jieba:
- dict.txt.small – 占用内存较小的词典文件
- dict.txt.big – 支持繁体分词更好的词典文件
The path is resolved from the project root directory.
jieba_dict_user¶
Use this setting to add a user dictionary to jieba without replacing the default dictionary. User dictionaries are useful for product names, technical terms, and other site-specific vocabulary:
The path is resolved from the project root directory.
Excluding sections and blocks¶
When Attribute Lists is enabled, a section can be excluded from the Lunr
index by adding data-search-exclude to its heading. The heading and all
content up to the next heading of the same or higher level are omitted:
# Page title
## Public section
This content is included.
## Internal section { data-search-exclude }
This complete section is excluded from Lunr.
Add the same attribute after an inline or block-level element to exclude only that element:
Metadata¶
boost¶
Use this property to increase or decrease the relevance of a page in Lunr
results. Values above 1 rank a page up and values below 1 rank it down:
exclude¶
Use this property to exclude a page and all of its subsections from the Lunr index. The Pagefind provider recognizes the same property: