From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
A round-trip study of tree-structured expression serialization in language models
Apple researchers propose a round-trip protocol to measure how much tree-structured compositional content survives when language models serialize expressions into natural language.
Evaluating 16 models, they find the channel is lossy and asymmetric, with at least 73.6% of failures originating at generation, and that fine-tuning on ∼3600 examples lifts open-weight models above an untrained frontier model.
From the source
Three main findings emerge. First, the channel is lossy and asymmetric: swapping which model generates and which extracts shifts accuracy by up to 60.4 points, and the best pair reaches 92.9% by combining different models on each end rather than the same model on both. Second, at least 73.6% of round-trip failures originate at generation, and difficulty is driven by tree structure (operator count, depth, right-branching) rather than model family. Third, the channel is trainable: ∼ 3600 fine-tuning examples that share the evaluation’s operators and tree shapes lift every open-weight model above untrained Gemini-3.1-Pro, an upper bound under matched semantics.
machinelearning.apple.com