fix(preview): parse admission tag attributes in order and count empty rows

XML allows a raw `>` and the other quote character inside an attribute
value, so a first-match search for `ref=`/`max=` could be fed a fake
value from an earlier attribute while saxes read the real one:

- <mergeCell>/<col> attributes are now read in order from the tag name
  with a sticky regex that consumes each quoted value whole. A tag whose
  attributes do not parse up to `>`, or that repeats a name, is refused.
- The chunk carry keeps everything from the last `<`, which can never
  appear inside an attribute value, instead of comparing against the
  last `>`.

ExcelJS keeps a Row object for every <row>, cells or not, so <row> tags
now count against per-sheet (100k) and total (250k) caps with their own
row-limit code, and the worker passes maxRows as a per-sheet backstop.

styles.xml counts every <xf> without tracking which list it sits in,
since a </cellXfs> inside a comment desynced that state.
This commit is contained in:
Aamer Akhter
2026-10-01 21:39:26 -04:00
parent e85b4f34dd
commit d65ee4f89d
6 changed files with 160 additions and 31 deletions
+6 -2
View File
@@ -127,8 +127,12 @@ async function loadWorkbook(bytes) {
const nextWorkbook = new self.ExcelJS.Workbook();
// ExcelJS expands every address of a `<dataValidation sqref>` into its own
// object (a whole-column dropdown is a million), and the preview never shows
// validations, so they are not parsed at all.
await nextWorkbook.xlsx.load(admitted, { ignoreNodes: ['dataValidations'] });
// validations, so they are not parsed at all. `maxRows` is a per-sheet
// backstop behind admission's row count, which also caps the workbook total.
await nextWorkbook.xlsx.load(admitted, {
ignoreNodes: ['dataValidations'],
maxRows: core.LIMITS.maxRowsPerSheet,
});
const nextSheets = new Map();
const nextRows = new Map();
normalizedStyles = [];