InductorParser 1.0.1067

dotnet add package InductorParser --version 1.0.1067
                    
NuGet\Install-Package InductorParser -Version 1.0.1067
                    
This command is intended to be used within the Package Manager Console in Visual Studio, as it uses the NuGet module's version of Install-Package.
<PackageReference Include="InductorParser" Version="1.0.1067" />
                    
For projects that support PackageReference, copy this XML node into the project file to reference the package.
<PackageVersion Include="InductorParser" Version="1.0.1067" />
                    
Directory.Packages.props
<PackageReference Include="InductorParser" />
                    
Project file
For projects that support Central Package Management (CPM), copy this XML node into the solution Directory.Packages.props file to version the package.
paket add InductorParser --version 1.0.1067
                    
#r "nuget: InductorParser, 1.0.1067"
                    
#r directive can be used in F# Interactive and Polyglot Notebooks. Copy this into the interactive tool or source code of the script to reference the package.
#:package InductorParser@1.0.1067
                    
#:package directive can be used in C# file-based apps starting in .NET 10 preview 4. Copy this into a .cs file before any lines of code to reference the package.
#addin nuget:?package=InductorParser&version=1.0.1067
                    
Install as a Cake Addin
#tool nuget:?package=InductorParser&version=1.0.1067
                    
Install as a Cake Tool

NuGet version

The Inductor Parser (IP) is a loose port of the Inductor C++ Parser. It's a PEG-style parser, designed for C#. Browse the full documentation site (guides, reference, and API) at https://ericzinda.github.io/InductorCSParser/. It supports and has been tested on .NET 8, 9 and 10, as well as Unity 6000.3.13f1 Standalone IL2CPP. What this means and how to test on other platforms is described in the Test Architecture Doc. It's released under the MIT License.

I ported this while creating a new project in Unity and during a period where I've been subjected to reviewing way too many Claude generated Regex's. I wrote up how I built this, if you're curious.

My goal is to design a parser library that is:

  • More Readable than Regex: The grammars are self-describing and human readable so they can be reasoned about, code reviewed and understood without looking up obscure letters and symbols.
  • Designed for World Languages: From the lexer, to the built-in rules, to normalization, it's designed around Unicode so grammars have a good starting point for world-language text (but it's not in your face if you don't care).
  • Safer Against Pathological Input: It's designed to avoid "catastrophic backtracking" and pitfalls like it that can hang your app or blow your stack, by default.
  • Able to run on WebGL and .NET Standard 2.1 (and later) using IL2CPP: It doesn't use Reflection.Emit or threads so that it can run in Unity targeting WebGL or IL2CPP on iPhone.
  • Fast enough to be used in production: It is competitive against other .Net Parsers and fast enough to be used as a regex replacement for most uses.
  • Easy to understand and customize: Your grammar is built out of simple rules that are easy to inspect and understand. Furthermore, building a new rule is simple and can do whatever you want: it is just code, not a grammar specific mathematical language

To use it from .NET, add the NuGet package:

dotnet add package InductorParser

For Unity, download InductorParser-<version>-netstandard2.1.zip from the latest release and put InductorParser.dll (and the .xml next to it, so IntelliSense shows the docs) in your project's Assets/Plugins folder. Both are built from the same source: the package has netstandard2.1 and net8.0 builds, and the zip is the netstandard2.1 one that Unity's IL2CPP backend loads.

If you just want to learn how to use it, follow the primers:

You can also just point Claude or Codex at it. I've used both rather interchangeably as tools when writing this parser and they both do a good job at understanding it, fixing bugs, describing how it works and how to use it, and using it directly in other projects.

For more background, read on.

More Readable than Regex

A major goal in building this parser is to replace hieroglyphic Regex patterns or complicated, hard to debug parsing code with something more readable, debuggable and understandable. Especially as I'm doing more and more reviewing of code written by LLMs, I've found it invaluable to have the LLM write pattern matching and parsing code in a form that I can actually review for correctness.

Compare a couple top Regex questions from StackOverflow:

Match numbers only (From https://stackoverflow.com/q/273141)

Regex: ^\d+$
Inductor Parser:

var numbersOnly = And(
    OneOrMore(OneOf(TokenSet.Digits)),
    Eof()
);

Match a line that doesn't contain the word "hede" (From https://stackoverflow.com/q/406230):

Regex: ^((?!hede).)*$
Inductor Parser (actually matches all end of line variants which the OP probably really wanted):

var lineWithoutHede = And(
    ZeroOrMore(And(
        Not(Literal("hede")),
        Not(EndOfLine()),
        AnyToken()
    )),
    EndOfLine(eofIsEol: true)
);

Primer: Building a Grammar walks through how to build rules in more detail.

Designed for World Languages

If you write grammars using the Inductor Parser, you get a foundation that supports Unicode from the start:

  • Each token presented to a rule is a user-perceived character (a "Grapheme Cluster" in Unicode) which keeps grammars from matching partial non-ASCII characters or emoji sequences accidentally and allows writing rules more naturally.
  • Built-in rules use Unicode-aware definitions for things like "whitespace" and "identifiers" so you don't miss corner cases.
  • The parser defaults to normalizing both the input and your rules to the same form (which you can choose) so that you can write rules how you want and they will match the different forms automatically.
  • Characters that can't possibly match the chosen normalized form for the input throw at compile time. They won't silently be ignored.
  • Every Symbol in the parse tree (and every error on the result) exposes its source position in four units: char index, token index, line, and column. These positions index into the original source even if it has been normalized into something else for parsing. Errors give you the single failure point the same way.

You can pretend you never heard the phrase "grapheme cluster" and write rules naturally: the guardrails are there by default and give you the right base to start from.

Here's a grammar for reading a simple setting that only accepts strings:

// Parse: Key = StringValue (e.g. Goo = 'some string')
var settingName = Identifier().As("name");

var quotedString = And(
    Token("'"),
    ScanUntil(Token("'")),
    Token("'"))
    .As("value");

var document = And(
    settingName,
    Optional(AnyWhitespace()),
    Token('='),
    Optional(AnyWhitespace()),
    quotedString
);

... and some examples that show how it handles different Unicode challenges well even if you weren't thinking about Unicode when you wrote it:

// Easy default case
var result = document.Parse("setting = '5'"); // name: "setting", value: "5"

// Identifier() follows official Unicode UAX #31 identifier rules, so names from many scripts work
document.Parse("Γειά = '5'");    // name: "Γειά",    value: "5"
document.Parse("привет = '1'");  // name: "привет",  value: "1"
document.Parse("你好 = '1'");    // name: "你好",     value: "1"

// é written as e + U+0301 (accent mark) is two C# chars that
// form one user-perceived grapheme. The parser accepts it
document.Parse("café = '5'");    // (é = e + U+0301) name: "café", value: "5"

// In a Devanagari language example, a grapheme can be a consonant
// joined to a virama or vowel sign, sometimes many C# chars long
document.Parse("नमस्ते = '1'"); // name: "नमस्ते", value: "1"

// 𠮷 is U+20BB7, one rune but two C# chars. 
document.Parse("𠮷田 = '5'"); // name: "𠮷田", value: "5"

// Optional(AnyWhitespace()) matches Unicode's White_Space property (UAX #44), not just ASCII
document.Parse("setting\u00A0=\u00A0'5'");  // (\u00A0 = non-breaking space)
document.Parse("setting\u3000=\u3000'5'");  // (\u3000 = ideographic space) name: "setting", value: "5"

// String values can hold anything except the closing quote. 
// Mixed scripts, emoji, and multi-rune graphemes all pass through untouched
document.Parse("motto = '你好 🎉 नमस्ते'"); // name: "motto", value: "你好 🎉 नमस्ते"
document.Parse("motto = '👨\u200D👩\u200D👧'");  // (ZWJ family emoji) name: "motto", value: "👨‍👩‍👧"
document.Parse("motto = '🇺🇸'");  // (regional-indicator flag) name: "motto", value: "🇺🇸"

// Emoji aren't in the UAX #31 identifier set, so the parser rejects them 
// the same way Python and Rust do:
document.Parse("setting🎉 = '5'"); // GrammarMismatch at char 7

Error positions are also designed for Unicode and reported in multiple units. When a letter takes more than one C# char, the char index and the token index diverge, and graphemes built from several joined characters (an emoji family, say) push them apart even further. This gives you the right tools for different jobs:

var result = document.Parse("𠮷田 = ");
// ErrorCharIndex=6, ErrorTokenIndex=5
// (𠮷 is one token but two chars, so char index runs one ahead; 田 is a normal one-char token)

var result = document.Parse("नमस्ते = ");
// ErrorCharIndex=9, ErrorTokenIndex=7
// (2 of the name's 4 letters take 2 chars each)

The same multi-unit positioning is available for every Symbol in the parse tree on success. Every Symbol has a SourceRange that exposes the same four fields (CharIndex, TokenIndex, Line, CharColumn) for both Start and End:

var result = document.Parse("motto = '👨‍👩‍👧'");
var range = result.Find(quotedString)!.SourceRange!.Value;
// Width of the matched value:
//   range.End.CharIndex  - range.Start.CharIndex  == 10  // 8 for the family + 2 quotes
//   range.End.TokenIndex - range.Start.TokenIndex ==  3  // 1 for the family + 2 quotes

Use whichever unit your code needs. Chars for string.Substring or an editor diagnostic. Tokens for a ^^^ underline a human will look at and recognize as covering one thing.

Primer: Unicode in the Inductor Parser walks through how Unicode works in rules in more detail.

Safer Against Pathological Input

Regex expressions can sometimes introduce denial-of-service attacks (or just plain poor user experiences) when they encounter adversarial or unexpected text. Here's a classic that looks reasonable in code review:

^([a-zA-Z0-9]+)*@example.com$

A simple email-ish validator. Feed it "aaaaaaaaaaaaaaaaaaaaaaaaa!" (25 a's) and .NET Regex will burn seconds trying to find a match. The problem is the nested + inside *: when the match fails, the engine has to try every way to split the a's across the two quantifiers before giving up. Add another a or two and the time doubles.

The Inductor Parser avoids this and is more readable as well:

var validator = And(
    OneOrMore(OneOf(TokenSet.Ascii.Letters | TokenSet.Ascii.Digits)),
    Literal("@example.com"),
    Eof()
);

OneOrMore greedily consumes all the a's in one pass, sees the !, fails cleanly. It takes linear time no matter what you throw at it.

Backtracking isn't the only way to hang. A 100 MB input file, a grammar that recurses 10,000 levels deep on nested parenthesis, or untrusted input in a web handler can all do it, too. The parser has three ways to handle these scenarios:

  • RuleCountLimit (default 10M) caps how much work a parse can do (rule invocations plus bulk-scan steps).
  • MaxDepth (default 1000) caps the recursion depth.
  • Timeout (default off) caps wall-clock time spent (done without a thread to support WebGL).

See Primer: Security-Related Concerns for more details on security related features and how the parser is designed to combat them.

Able to run on WebGL and .NET Standard 2.1 (and later) using IL2CPP

Inductor Parser is designed to be able to be used in Unity, targeting WebGL and iPhone, which constrains it:

  • WebGL is single-threaded, so no background timers
  • IL2CPP means no IL can be generated at runtime: No System.Reflection.Emit, no LINQ Expression.Compile, no source generators producing IL at parse time
  • .NET Standard 2.1, not .NET 5+ since Unity's IL2CPP surface is still netstandard2.1. (works fine on .NET 5+, though!)
  • Unity's runtime splits text into characters using outdated Unicode rules (emoji sequences and CRLF come apart), and its string.Normalize misses conversions .NET applies, so the build Unity loads ships its own copy of .NET's character segmentation and its own UAX #15 normalizer and uses both automatically. Other runtimes can opt into the same implementations at startup using UnicodeEnvironment.Implementation when they need parse trees identical to a Unity client's. UnicodeGotchas.md has the details.

Fast Enough to be Used in Production

To evaluate performance I used open source benchmarks built by others so that I wasn't unfairly building tests that IP was good at. You can run them yourself in the src/Benchmarks folder.

The Parlot project had a great benchmark of C# parser libraries that I forked into the src/Benchmarks folder. I added both InductorParser and Pegasus (another PEG-style parser) to the suite. You can read the details of the test, what I changed, etc here. It asks each parser library to build a Json parser and read 4 different documents that are different shapes. Real world and a nice benchmark. In addition to performance, it's illustrative to look a the grammars for each parser library and compare for readability and reviewability, they're here.

Latest results

Performance chart

Both the JPG above and the interactive HTML version are regenerated automatically every benchmark run, with the run date stamped in the chart title so you can tell at a glance how fresh the numbers are. The full table with allocations and ratios is in src/Benchmarks/README.md.

Product Compatible and additional computed target framework versions.
.NET net5.0 was computed.  net5.0-windows was computed.  net6.0 was computed.  net6.0-android was computed.  net6.0-ios was computed.  net6.0-maccatalyst was computed.  net6.0-macos was computed.  net6.0-tvos was computed.  net6.0-windows was computed.  net7.0 was computed.  net7.0-android was computed.  net7.0-ios was computed.  net7.0-maccatalyst was computed.  net7.0-macos was computed.  net7.0-tvos was computed.  net7.0-windows was computed.  net8.0 is compatible.  net8.0-android was computed.  net8.0-browser was computed.  net8.0-ios was computed.  net8.0-maccatalyst was computed.  net8.0-macos was computed.  net8.0-tvos was computed.  net8.0-windows was computed.  net9.0 was computed.  net9.0-android was computed.  net9.0-browser was computed.  net9.0-ios was computed.  net9.0-maccatalyst was computed.  net9.0-macos was computed.  net9.0-tvos was computed.  net9.0-windows was computed.  net10.0 was computed.  net10.0-android was computed.  net10.0-browser was computed.  net10.0-ios was computed.  net10.0-maccatalyst was computed.  net10.0-macos was computed.  net10.0-tvos was computed.  net10.0-windows was computed. 
.NET Core netcoreapp3.0 was computed.  netcoreapp3.1 was computed. 
.NET Standard netstandard2.1 is compatible. 
MonoAndroid monoandroid was computed. 
MonoMac monomac was computed. 
MonoTouch monotouch was computed. 
Tizen tizen60 was computed. 
Xamarin.iOS xamarinios was computed. 
Xamarin.Mac xamarinmac was computed. 
Xamarin.TVOS xamarintvos was computed. 
Xamarin.WatchOS xamarinwatchos was computed. 
Compatible target framework(s)
Included target framework(s) (in package)
Learn more about Target Frameworks and .NET Standard.
  • .NETStandard 2.1

    • No dependencies.
  • net8.0

    • No dependencies.

GitHub repositories

This package is not used by any popular GitHub repositories.

Version Downloads Last Updated
1.0.1067 0 10/5/2026