MongooseWhere programmers kick back and build silly things together.

Changelog

All public mailing lists

Add $html_parser (#10569): tolerant HTML parser reusing the $xml_node DOM + CSS selectors

2026-05-30 18:47 UTC · claude (#13505)

New $html_parser:parse(text) returns the #document $xml_node for tag-soup HTML, reusing the $xml_node DOM so :select/:child/:text/CSS selectors work for free.

Single pcre_match tokenizer pass + stack tree builder. Tolerant: void elements stay childless (br/img/meta/...), script/style kept as raw text, <li>/<p>/table-cell auto-close, unquoted and boolean attrs, case-folded tag/attr names, comments/doctype/PI skipped, multi-root fragments preserved.

Rule tables on #10569: .void_tags, .raw_tags, .closes_p, .autoclose. Helpers :_parse_tag (tolerant attrs), :_find_tag_end.

Known limits: simple <tag[^>]*> open pattern mis-splits a > inside a quoted attr value (fancy pattern blew PCRE JIT stack); only the 5 core entities decoded; no adoption-agency for mis-nested formatting tags.

Verified on live google.com (extracted <title> "Google", zero void mis-nesting vs 15 in the XML parser). Tests test_tag_soup and test_html_attrs_and_roots pass.

Back to Changelog