Introduction to PseudoLangs
Writing prompts as structured pseudo-code instead of plain English can measurably improve what models get right. Research on prompt-format sensitivity found that meaning-preserving formatting changes alone swing accuracy, and one study that rewrote 132 tasks as pseudo-code beat natural-language prompts by a wide margin. Compressing prompts by dropping filler tokens works but costs a small amount of accuracy, while symbol-based encoding mostly fails because emojis split into multiple tokens and read unreliably. The honest rule: structure is legitimate when the model actually follows it, not when it's decorative.
Notes
PseudoLangs: Family overview (Prompt Engineering Institute, feed, 2026-06-25)
Constructed prompt notations between free-form prose and raw code. Claims format encoding — not just content — materially changes model results: "meaning-preserving changes to formatting alone can swing a model's accuracy by a large margin."
The spectrum (dial from natural language):
- PseudoScript/SudoCode — structured pseudo-code: numbered steps, named functions, variables, explicit control flow → precision.
- MiniScript — minified prompting that drops filler to fit token budget → efficiency.
- SymboScript — symbolic glyph encoding → density.
Evidence per member:
- PseudoScript (validated, production-ready): Mishra et al. rewrote 132 tasks as pseudo-code (Pseudo-Code Instructions), beating natural-language prompts "by a wide margin" via reduced prose ambiguity. Corroborated by Code Prompting Elicits Conditional Reasoning, which credits improved variable-state tracking.
- MiniScript (real, best automated): LLMLingua compresses prompts by dropping low-information tokens for "a small, measurable accuracy cost."
- SymboScript (experimental frontier): the intuition that symbols are free compression is "mostly wrong" — most emojis tokenize into multiple sub-tokens and models read uncommon symbols unreliably. Principled version: learn compact tokens (gist-token compression), not glyphs bolted onto inference.
Line that keeps it honest: a notation is legitimate "when the model acts on it and a human can maintain it"; it's theatre when decorative — syntax describing a process the model never executes, layered on to look rigorous. Cut notation only you perform.
Says the following articles take each member in turn.
Now done — let me mark the task complete.
Done. Notes above (~170 words) captured the spectrum, per-member evidence with study names, and the caveat that symbolic encoding mostly doesn't compress.