← David Mashiah

Research

Published work, with DOIs

Everything below is deposited, versioned and citable. Each entry links to the record itself rather than to a description of it, and the negative results are here alongside the rest — a benchmark that had to be rebuilt after its own audit found a flaw is on this page for the same reason the others are.

ORCID 0009-0004-4684-955X

  1. Journal contribution · 2 July 2026

    Algebraic Encoding of Amino Acid Sequences

    Algebraic Encoding of Amino Acid Sequences: Extending Finite Field Models to GF(32)

    Extends finite-field encoding of amino acid sequences to GF(32): a field sized by the alphabet rather than inherited from a machine word, so that a twenty-letter alphabet fits in one symbol with room to spare. Byte-sized arithmetic leaves most of the field unused on an alphabet this small, which makes every parity symbol cost far more than the information it protects. This is the encoding the Reed–Solomon toolkit deposited alongside it is built on.

    • Algebraic coding theory
    • Finite field encoding
    • Amino acid sequences
    doi.org/10.6084/m9.figshare.32885588
  2. Software · 11 July 2026

    GF(25) Encoding & Reed–Solomon Error Correction

    A Reproducible Toolkit for GF(2^5) Encoding and Reed–Solomon Error Correction of Synthetic Biopolymer Data

    A reproducible toolkit implementing Reed–Solomon error correction over GF(2^5) for synthetic biopolymer data — the same construction that repairs a damaged barcode or QR symbol instead of only detecting the damage. The field is sized by the sequence alphabet, four nucleotides or twenty amino acids, rather than by the byte a barcode inherits. Reed–Solomon codes are maximum distance separable, so the minimum distance is exactly n − k + 1, which is the best any code achieves for that overhead.

    • Reed–Solomon error correction
    • Finite field encoding
    • Error-correcting codes
    doi.org/10.6084/m9.figshare.32963456
  3. Dataset · 15 July 2026

    MetaXFam

    MetaXFam: A Cross-Family Dataset and Toolkit for Mechanical Metamaterial Homogenization Surrogates

    An open dataset and toolkit for studying how machine-learning surrogates of mechanical metamaterial properties behave on unit-cell topologies they were not trained on. Includes a from-scratch 2D numerical homogenization engine and a validation suite checked against exact analytical results.

    • Mechanical metamaterials
    • Numerical homogenization
    • Machine learning reliability
    • Out-of-distribution generalization
    doi.org/10.6084/m9.figshare.32993597
  4. Preprint · under review · 24 July 2026

    Perfect In-Distribution Accuracy Does Not Imply Learned Physics

    Perfect In-Distribution Accuracy Does Not Imply Learned Physics: Measuring Cross-Family Generalization of Metamaterial Homogenization Surrogates

    A preprint measuring whether machine-learning surrogates of metamaterial homogenization learn transferable structure or only fit the families they were trained on. In-distribution accuracy is shown not to predict cross-family performance. Under review at Extreme Mechanics Letters; the benchmark it reports was deposited as MetaXfam22 and has since been rebuilt and renamed MetaXFam-D.

    • Mechanical metamaterials
    • Numerical homogenization
    • Machine learning reliability
    • Out-of-distribution generalization
    doi.org/10.5281/zenodo.21534865
  5. Dataset · 29 August 2026

    MetaXFam-D

    MetaXFam-D: a distinct-rasterisation benchmark for cross-topology generalization of metamaterial homogenization surrogates

    A benchmark for cross-topology generalization of metamaterial homogenization surrogates, rebuilt so that no two cells share a geometry. The version before it sampled continuous shape parameters onto a 48×48 grid, where many parameter values collapse to the identical pixel image: 61.6 per cent of its cells were duplicates, and under a random 75/25 split 67.7 per cent of test cells had their exact geometry in the training set, so its in-distribution and random-split scores are inflated by memorisation. Cross-family results are much less affected. This version ships 18 families of 172 cells verified distinct, an expanded pool covering all 29 families, the generators, a validated periodic homogenization solver and the duplication-audit tools.

    • Mechanical metamaterials
    • Numerical homogenization
    • Machine learning reliability
    • Out-of-distribution generalization
    • Training set selection
    doi.org/10.5281/zenodo.21597150
  6. Software · with preprint · 31 August 2026

    Sub-cell Eigenmode Descriptors

    Sub-cell eigenmode descriptors for metamaterial homogenisation surrogates: code, descriptors, results and preprint

    Code, descriptors, complete results and a venue-neutral preprint for a study of why machine-learning surrogates of the homogenized stiffness tensor fail to transfer to unseen unit-cell topologies. Version 2 recomputes everything on duplicate-free data and, in the record's own words, removes the central claim of version 1: the three-number descriptor's advantage over a 2,304-pixel baseline falls from +0.364 at a 100 per cent win rate to +0.045 under a random forest and is absent on C11, while the pixel baseline itself recovers from -0.008 to +0.37 on C22. What survives is stability rather than magnitude — held-out R2 spans +0.02 to +0.41 for the descriptor against -42.99 to +0.34 for raw pixels — and the size of the evaluation gap: the same surrogate scores 0.90 to 0.96 under a random split and -0.73 to +0.14 when whole topology families are held out. Version 1 is superseded but not withdrawn.

    • Mechanical metamaterials
    • Numerical homogenization
    • Machine learning reliability
    • Out-of-distribution generalization
    • Micropolar elasticity
    doi.org/10.5281/zenodo.22137052