Agent Skills: Attention Variants from Papers

Implement paper-defined attention variants by extracting mechanism invariants, mapping tensor shapes, preserving module interfaces, validating numerical behavior, and integrating the custom attention into an existing transformer stack.

UncategorizedID: benchflow-ai/skillsbench/attention-variants-from-papers

Repository

benchflow-aiLicense: Apache-2.0
1,819369

Install this agent skill to your local

pnpm dlx add-skill https://github.com/benchflow-ai/skillsbench/tree/HEAD/tasks-extra/diff-transformer_impl/environment/skills/attention-variants-from-papers

Skill Files

Browse the full folder contents for attention-variants-from-papers.

Download Skill

Loading file tree…

tasks-extra/diff-transformer_impl/environment/skills/attention-variants-from-papers/SKILL.md

Skill Metadata

Name
attention-variants-from-papers
Description
Implement paper-defined attention variants by extracting mechanism invariants, mapping tensor shapes, preserving module interfaces, validating numerical behavior, and integrating the custom attention into an existing transformer stack.

Attention Variants from Papers

Use this skill when a paper changes how attention scores, branches, normalization, or head sharing work, but the module still needs to behave like a drop-in transformer attention block.

Workflow

  1. Read the paper for invariants, not names. Use paper-to-implementation.md to extract the external contract, the changed computation, and the training-time constraints.
  2. Build a shape ledger before coding. Use shape-ledger.md to track projections, head grouping, branch count, and output width.
  3. Choose the mechanism pattern. Use mechanism-patterns.md for subtractive attention, branch mixing, learned gates, and extra normalization.
  4. Preserve the module boundary. Keep the same input and output shape, mask semantics, positional encoding flow, and cache behavior unless the task explicitly changes them.
  5. Validate in layers. Start with random-tensor smoke tests, then compare against a baseline attention path. Use stability-and-validation.md.
  6. Integrate into the stack last. Swap the new module into one transformer block, verify the residual path, then roll it through the full model. Use transformer-integration.md.

Checklist

  • extract the paper's invariants before writing code
  • account for every reshape, branch, and repeat in a shape ledger
  • preserve output width at concatenation or output projection
  • apply masks and positional terms at the intended stage
  • confirm random smoke tests stay finite
  • compare unchanged behaviors against a baseline attention implementation

Reference Map