skillstested.com
Skill paths / Landing page → visual QA
Evidence-backed skill path · v1

Landing page audit, implementation patch, and final OG

PARTIAL EVIDENCE

Use hallmark for judgment, kill-ai-slop for mechanical recall, and visual-skills for social-card direction, then deliver an inspectable page patch and final OG image.

OUTCOMEA prioritized design audit, raw and human-reviewed scanner report, applied page patch, and final 1200×630 OG PNG.

Bring

  • Public URL, local page, or repository
  • Page purpose and intended audience
  • Declared design genre or design-system constraints

Receive

  • Severity-ranked visual audit
  • Scanner report with human-confirmed findings
  • Applied static-page patch and correction checklist
  • Final 1200×630 OG PNG

The path and every handoff

  1. Audit with judgment hallmark
    FUNCTION VERIFIED

    Find hierarchy, genre, typography, and macrostructure problems that mechanical rules cannot judge.

    Input
    Rendered page plus its audience, purpose, and declared design genre
    Instruction
    Use hallmark's audit verb and record each finding as what, where, severity, and proposed fix without editing the page yet.
    Output
    hallmark-audit.md human confirmation
  2. Scan and confirm anti-patterns kill-ai-slop
    FUNCTION VERIFIED

    Catch mechanical regressions quickly, then separate true problems from design-system exemptions and scanner self-matches.

    Input
    Page source and declared design-system exemptions
    Instruction
    Run the bundled scanner, retain raw hits, and complete its required human-confirm step before proposing fixes.
    Output
    anti-slop-report.md human confirmation
  3. Create the social card direction visual-skills (video + image)
    FUNCTION VERIFIED

    Choose an image model based on the content and convert the approved page direction into a constrained OG prompt.

    Input
    Approved brand direction, exact copy, palette, and 1200×630 use case
    Instruction
    Follow visual-skills' mandatory model-routing reference, freeze the exact copy and safe margins, then turn the approved direction into a final vector-first 1200×630 asset.
    Output
    og-image-direction.md human confirmation
  4. Verify the corrected page and OG asset Verification or human checkpoint
    WORKFLOW VERIFIED

    Check that fixes did not introduce regressions and that the final image meets its dimensions and copy contract.

    Input
    Corrected page and generated OG image
    Instruction
    Re-run mechanical and structural checks, confirm the analytics and privacy scripts are byte-preserved, and inspect OG dimensions, text, and safe margins.
    Output
    before-after report and final-og.png human confirmation

Result you can inspect now

EVIDENCE REPLAY Three design skills applied to the same live homepage
2026-07-25
Open the run record — what it proves, what it doesn't, and the raw artifacts

What this proves: hallmark produced a severity-ranked audit, kill-ai-slop ran its scanner and human-confirm pass, and visual-skills produced a model-routed 1200 by 630 social-card brief for the same page.

What it does not prove: Applying the proposed fixes, generating a final OG image, repeated-run consistency, or measuring quality gain against a no-skill baseline.

  • hallmark: 0 critical, 1 major, 3 minor findings before genre confirmation
  • kill-ai-slop: 30 raw hits across 3 groups, about 1 surviving human confirmation
  • visual-skills: one structured GPT Image social-card brief

Open the full source record →

EVIDENCE REPLAY Nine independent landing-page delivery runs with final OG assets
2026-08-02
Open the run record — what it proves, what it doesn't, and the raw artifacts

What this proves: The same production homepage was run three times with no skill context, Hallmark alone, and the published three-skill combination. All nine runs produced a revised full HTML page, substantive audit, and renderable 1200 by 630 OG; all nine preserved the analytics and goal-privacy scripts byte for byte and passed the same structural delivery checks.

What it does not prove: That the skill combination is higher quality than the control, that arbitrary customer frameworks need no rework, or that a paid order can always be fulfilled inside three business days. Scanner medians were identical across arms.

  • Valid complete deliveries: 9 of 9
  • Median duration: control 126.30s, single skill 229.87s, combination 168.14s
  • Median scanner result in every arm: 4 groups and 26 raw hits
  • Selected final asset: combination run 2, chosen for clearest product-path communication rather than a claimed arm-level quality win
SkillsTested social card showing outcome, path, and evidence as a three-step flow
Selected final OG image
Nine-run mechanical summary →

Known limits

Next verification required: Fulfil the first non-owner order within the promised three-business-day window and record rework, refund, and buyer acceptance without publishing private inputs.

Implementation pack

The pack is generated from this path record and the tested-skill source data. Install commands, commits, evidence status, handoffs, and limitations stay attached to the plan.