- JavaScript 76.7%
- Go 19.9%
- templ 1.7%
- CSS 1.6%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
|
All checks were successful
ci/woodpecker/push/change_log Pipeline was successful
ci/woodpecker/tag/release Pipeline was successful
ci/woodpecker/push/ci Pipeline was successful
ci/woodpecker/cron/vulnerability Pipeline was successful
ci/woodpecker/cron/renovate Pipeline was successful
ci/woodpecker/cron/release Pipeline was successful
ci/woodpecker/cron/auto_merge Pipeline was successful
Reviewed-on: #8 |
||
| internal | ||
| test | ||
| .gitignore | ||
| .goreleaser.yaml | ||
| .markdownlint.json | ||
| .prettierrc.json | ||
| .sqruff | ||
| CLAUDE.md | ||
| Containerfile | ||
| go.mod | ||
| go.sum | ||
| main.go | ||
| README.md | ||
| renovate.json | ||
| sqlc.yaml | ||
srchfor
srchfor is a private web app to read a tree of documentation and to search it. It indexes the English text of every page with Bleve, shows the pages next to their table of contents, and offers the same actions through a REST API.
Features
- Single, statically linked Go binary. No external services: the only stores are a SQLite database and a Bleve index on local disk.
- Reads Markdown, HTML and plain text pages. Links between pages and images are rewritten to work inside the app, and all HTML is sanitized.
- Full-text search with English stemming, "exact phrases", highlighted snippets, product, version and book filters and facets. Matches in a title rank above matches in a heading, which rank above matches in the body. Words with typos are found, a half-typed last word finds longer words, and "did you mean" suggests the right spelling.
- A hit opens the page at the first match, with the searched words marked and buttons to step through the matches, which helps on pages that are hundreds of kilobytes long.
- Navigation from the
toc*.jsonfiles of each book, shown as a compact slice around the current page so even books with ten thousand pages stay fast. - Incremental indexing: only added, changed or removed files are read again. The index is brought up to date at startup, whenever the documentation directory changes (a file watcher), on request from the admin page or the API, and optionally on a timer.
- Operations: a health report with one check per component at
/healthz, and Prometheus metrics on a separate listen address. - Server rendered HTMX frontend with the same "paper and highlighter" look as Learn IT.
- A REST API under
/api/, built contract first from an OpenAPI 3.1 spec, with a hosted Swagger UI at/api/docs/. - The site is private. Login goes through an external OIDC provider (for example Authelia) or a mock login page for local development. People without an account can get in with an access code an admin creates. The API is authenticated with Personal Access Tokens.
Documentation tree
SRCHFOR_DOCS_DIR points at a directory laid out as <product>/<version>/<page>:
docs/
zos/
3.2.0/
sdsf.md
sdsf-action-characters.md
toc-sdsf.json
assets/9e92e8dda9506e78.gif
- Pages are
.md,.htmlor.txtfiles. A Markdown page may start with front matter holdingtitle,source(the original URL) andfetched; without a title the first heading is used. - Links such as
[x](other-page.md#part)and images such asare relative to the page. Theassetsdirectory is never indexed as pages and only image files are served from it. toc.json(ortoc-<book>.json) holds a book's navigation as[[slug, title, depth], ...]in reading order, whereslugis a page's path below the version directory without extension. Pages that no table of contents lists are still readable and searchable.- The directory is only read, never written.
Requirements
- Go (see
go.modfor the exact version). templ,sqlc,oapi-codegenandgolangci-lintfor development, for example installed withmise use -g templ@latest.
Building
go generate ./...
templ generate
CGO_ENABLED=0 go build .
The build must run with CGO_ENABLED=0 to keep the binary statically linked. You can check this
afterward with file srchfor (it should say "statically linked").
Running
The binary takes a subcommand:
srchfor servestarts the web server and the background index run.srchfor checkcalls the running server's/healthzendpoint and exits 0 or 1 (see below). This is meant to be used as a container health check.
All configuration is supplied through environment variables, each prefixed with SRCHFOR_:
| Variable | Description |
|---|---|
SRCHFOR_DATA_DIR |
Directory for the SQLite database and the search index. Required. |
SRCHFOR_DOCS_DIR |
The documentation tree, see above. Required. |
SRCHFOR_LISTEN_ADDR |
Address to listen on. Defaults to :8080. |
SRCHFOR_RESCAN_INTERVAL |
Run an incremental index this often, for example 6h. Off by default. |
SRCHFOR_WATCH |
Watch the documentation directory for changes. true by default. |
SRCHFOR_WATCH_DELAY |
Quiet time after a change before indexing, at least 1s. Default 5s. |
SRCHFOR_METRICS_ADDR |
Serve Prometheus metrics on this address, for example :9100. Off by default. |
SRCHFOR_AUTH_MODE |
Either oidc or mock. Required. |
SRCHFOR_MOCK_AUTH_FILE |
Path to a fixture file of mock users, used when SRCHFOR_AUTH_MODE=mock. |
SRCHFOR_OIDC_ISSUER_URL |
Base URL of the OIDC provider, used when SRCHFOR_AUTH_MODE=oidc. |
SRCHFOR_OIDC_CLIENT_ID |
OIDC client ID. |
SRCHFOR_OIDC_CLIENT_SECRET |
OIDC client secret. |
SRCHFOR_OIDC_REDIRECT_URL |
The app's own /oidc/callback URL. |
For local development, source test/.env or set SRCHFOR_AUTH_MODE=mock and
SRCHFOR_MOCK_AUTH_FILE=test/mockauth.json. This gives you a login page listing a few fixture users
(an admin and a reader), so you do not need a real identity provider:
SRCHFOR_DATA_DIR=/tmp/srchfor SRCHFOR_DOCS_DIR=./docs SRCHFOR_LISTEN_ADDR=:9000 \
SRCHFOR_AUTH_MODE=mock SRCHFOR_MOCK_AUTH_FILE=test/mockauth.json ./srchfor serve
The mock login sets a Secure cookie, which browsers accept for http://localhost.
Operations
Keeping the index current
The index is brought up to date at startup. After that, a file watcher (SRCHFOR_WATCH, on by
default) notices added, edited, removed and renamed documents and toc*.json files, waits until
nothing has changed for SRCHFOR_WATCH_DELAY, and then starts an incremental run, which only reads
what changed. A scraper or a copy that writes thousands of files therefore causes one run, not
thousands. Changes in the assets directories are ignored. The runs show up on the Index page with
the trigger watch.
Some file systems send no change events (certain network mounts), and the system limit of watched
directories can be reached. The server then keeps working, the health report shows the watcher as
degraded, and you can switch the watcher off with SRCHFOR_WATCH=false and use
SRCHFOR_RESCAN_INTERVAL instead.
Health
GET /healthz returns JSON with one check per component and needs no login:
{
"status": "ok",
"version": "1.2.0",
"uptimeSeconds": 4021,
"checks": {
"database": { "status": "ok" },
"index": { "status": "ok" },
"indexing": { "status": "ok" },
"docs": { "status": "ok" },
"watcher": { "status": "ok" }
}
}
statusisok,degraded(the service works, but look at it) orunhealthy. The HTTP status is 200 forokanddegradedand 503 forunhealthy, which is whatsrchfor checkand a container health check test.database,indexanddocs(the documentation directory) arefailed, and the serviceunhealthy, when they cannot be used.indexisdegradedwhen its document count differs from the records of what was indexed, which the next run repairs.indexingisdegradedwhen the last finished run failed or a run takes more than 30 minutes.watcherisdegradedwhen it was asked for but could not start.
Messages are short and contain no paths or document counts; those are in the metrics.
Metrics
Set SRCHFOR_METRICS_ADDR (for example :9100) to serve Prometheus metrics at /metrics on that
address. They are never served on the main port, and they have no authentication: bind the
address to an internal network (for example 127.0.0.1:9100, or a port that only your Prometheus
can reach). Without the variable there is no metrics endpoint.
scrape_configs:
- job_name: srchfor
static_configs:
- targets: ["srchfor.internal:9100"]
Besides the Go runtime and process metrics, all prefixed srchfor_:
| Metric | What it tells you |
|---|---|
http_requests_total{route,method,code} |
Requests per route pattern, such as GET /read/{product}/{version}/{slug...}. |
http_request_duration_seconds{route,method} |
Request latency histogram. http_requests_in_flight is the current load. |
search_requests_total{surface,outcome} |
Searches by web or api and by outcome: hits, none, approximate, unavailable, error. |
search_duration_seconds{surface}, search_suggestions_total |
Search latency, and how often a "did you mean" was offered. |
index_documents, index_size_bytes |
Documents in the index and its size on disk. |
index_documents_by_version{product,version} |
Indexed documents per product version. |
index_running, index_runs_total{mode,trigger,status} |
Whether a run is active, and finished runs by mode, trigger and result. |
index_run_duration_seconds |
Duration histogram of finished runs. |
index_last_run_timestamp_seconds, _duration_seconds, _success, _documents{kind} |
The last finished run: when, how long, whether it worked, and added, updated, removed, unchanged and failed documents. |
logins_total{provider,result} |
Logins by oidc or mock and ok, denied or error. |
guest_redemptions_total{result}, guest_codes{state} |
Guest code attempts (ok, invalid, throttled) and codes by state. |
sessions_active |
Sessions that have not expired. |
api_auth_failures_total{reason} |
Rejected API requests: missing, invalid, no-role. |
watcher_enabled, watcher_events_total, watcher_triggered_runs_total |
The file watcher. |
build_info{version,goversion} |
The running version. |
No label ever holds a query, a path, a page or a user, so the number of series stays small. Useful
alerts: srchfor_index_last_run_success == 0, time() - srchfor_index_last_run_timestamp_seconds
growing while SRCHFOR_RESCAN_INTERVAL or the watcher is on, a rising share of
search_requests_total{outcome="none"}, and up == 0.
Access
Everything except the login pages, the access code page and static files needs a login. Two groups of the identity provider decide the role, and a user in neither is denied:
srchfor_membercan read and search.srchfor_admincan do the same, and also start index runs, create Personal Access Tokens and manage guest access codes.
Groups are never edited in the app. They are read from the provider at each login.
Guest access codes
An admin can let people without an account in on the Guests page (/admin/guests): give a code
a label such as "Acme team", choose an expiry date or "never expires", and press "Copy link". The
link looks like https://your-host/guest?code=XK7P9-QRT4M-2K8N3-W1H6D. Opening it, or typing the
code at /guest, signs the visitor in as a guest for 30 days (never longer than the code lives),
remembered in an HttpOnly cookie.
- A code is shared and reusable. Give it to as many people as you like; nothing records who used it.
- A guest can read and search everything, but has no profile, no tokens, no admin pages and no API access.
- Revoking a code (or its expiry) ends all its guests' sessions on their very next request.
- Wrong codes are throttled: after 20 failed attempts within a minute from one address, further attempts are refused for a while. Behind a reverse proxy every visitor shares the proxy's address, so the throttle then applies to all of them together.
- The login screen (
/login, where every visitor without a session lands) has a "Log in" button for the identity provider and a link to/guest, so a guest can also just type the code there. Logging out returns to this screen. Note that the identity provider keeps its own single sign-on session, so pressing "Log in" again signs you straight back in without asking for a password.
REST API
The API is described by internal/api/openapi.yaml, and the running server hosts a Swagger UI at
/api/docs/. Every endpoint needs Authorization: Bearer <token>. An admin creates a token on the
profile page (or with POST /api/tokens); the value is shown once, expires on the chosen date, and
can be revoked. A token acts with the role its owner had at their latest login, so a demoted admin's
tokens lose admin actions, and a user without any role loses them altogether. Admins can also manage
guest codes with GET/POST /api/guest-codes and DELETE /api/guest-codes/{id}.
curl -H "Authorization: Bearer $TOKEN" "http://localhost:9000/api/search?q=opercmds&limit=5"
curl -H "Authorization: Bearer $TOKEN" "http://localhost:9000/api/pages/zos/3.2.0/sdsf?format=raw"
curl -H "Authorization: Bearer $TOKEN" -X POST -d '{"mode":"full"}' \
-H "Content-Type: application/json" http://localhost:9000/api/index/runs
Search syntax
Every word must occur in the title, a heading or the body of a page. English stemming applies
(widget finds widgets) and very common words (the, up) are ignored. Names such as
JES2MON.DISPLAY.DETAIL are split at dots, slashes and colons, so each part is searchable. Put
words in double quotes to find them together and in order.
A match in the title weighs most, then one in a heading, then one in the body, so a page with the
word in a heading comes before one that only mentions it. A heading that matched is shown above the
snippet, marked with a §.
Typos are tolerated. When fewer than 20 pages match exactly, words of four or more letters also find similar words: one wrong, missing or extra letter up to nine letters, two from ten letters, as long as the first letter is right. Words containing digits and quoted phrases always match exactly, and exact matches rank above similar ones. When nothing matches exactly, the page says so.
The word you are still typing, the last one, also finds longer words: sdsf pan finds panels, and
JES2MON.DISP finds JES2MON.DISPLAY.DETAIL. It needs three letters and works only while the query
does not end in a space or a closing quote; it applies in the same cases as typo tolerance.
When a word looks misspelled, the page says "Did you mean …?" with the most common word of the
documentation within one or two edits (autmation becomes automation), if that finds more pages.
The same suggestion is in the API as suggestion.
Opening a hit takes you to the page with the searched words marked (the hl parameter of the link)
and scrolled to the first match, preferring one in a heading. The bar above the page has buttons for
the previous and next match and a link that removes the marks.
Deployment
The project ships a Containerfile and is meant to run as a container. On a systemd based host with
Podman, it can be run as a Quadlet unit, for example:
[Unit]
Description=srchfor service
StartLimitBurst=5
StartLimitIntervalSec=90
[Container]
Image=localhost/srchfor:latest
# Storage options
Volume=srchfor.volume:/data
Volume=/srv/docs:/docs:ro
# Network options
PublishPort=127.0.0.1:9000:9000
# Metrics have no authentication: publish them to localhost or an internal network only
PublishPort=127.0.0.1:9100:9100
# Environment options
Environment="SRCHFOR_OIDC_CLIENT_ID=<your-oidc-client-id>"
Environment="SRCHFOR_OIDC_CLIENT_SECRET=<your-oidc-client-secret>"
Environment="SRCHFOR_OIDC_ISSUER_URL=https://your-authelia-host"
Environment="SRCHFOR_METRICS_ADDR=:9100"
# Healthcheck options
HealthCmd=[ "srchfor", "check" ]
HealthInterval=30s
HealthRetries=10
HealthStartPeriod=15s
HealthTimeout=15s
[Service]
Restart=on-failure
RestartSec=2
The /data volume holds the database and the index between restarts, /docs is the documentation
tree. Replace the OIDC values with your own.
CLAUDE.md in the repository root has the architecture and the styling guide.
License
GPL 3.0.