commit a9b7d78cead24f5d291f96529933395c89671e5c
parent da200169ee7fd530f595063d13675ac3b271b87e
Author: Achuthan TM <achuthantm05@gmail.com>
Date: Mon, 23 Mar 2026 16:51:25 +0530
Week 1: Add gitignore and detailed architecture docs (Part 1: Initial definitions)
Diffstat:
| A | .gitignore | | | 18 | ++++++++++++++++++ |
| A | docs/details.md | | | 84 | +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ |
2 files changed, 102 insertions(+), 0 deletions(-)
diff --git a/.gitignore b/.gitignore
@@ -0,0 +1,18 @@
+# Prerequisites
+*.d
+
+# Compiled Object files
+*.slo
+*.lo
+*.o
+*.obj
+
+# Precompiled Headers
+*.gch
+*.pch
+
+# Linker files
+*.ilk
+
+# Debugger Files
+*.pdb
diff --git a/docs/details.md b/docs/details.md
@@ -0,0 +1,84 @@
+# Stage 1 — C++ Regex Parser
+
+This stage is responsible for transforming a raw regular expression string into an Abstract Syntax Tree (AST) that can be processed by the NFA construction stage.
+
+---
+
+## 1. Lexer Details
+
+### Overview
+
+The lexer is the first component. Its primary responsibility is to read a raw regular expression string character-by-character and transform it into a sequence of typed tokens. This abstraction simplifies the downstream parsing process by handling character-level concerns like literal character validation and metacharacter identification.
+
+### Token Types
+
+The lexer identifies the following token types:
+
+| Token Type | Representation | Description |
+| -------------- | ---------------------------- | --------------------------------------------------- |
+| `LITERAL` | Any printable ASCII (32–126) | Represents a literal character match. |
+| `DOT` | `.` | Matches any single character. |
+| `STAR` | `*` | Kleene star: zero or more repetitions. |
+| `PLUS` | `+` | One or more repetitions. |
+| `QUESTION` | `?` | Optional: zero or one occurrence. |
+| `PIPE` | `\|` | Union / Alternation operator. |
+| `LPAREN` | `(` | Start of a subexpression group. |
+| `RPAREN` | `)` | End of a subexpression group. |
+| `END_OF_INPUT` | `\0` | Sentinel token marking the end of the regex string. |
+
+### Input Processing Rules
+
+1. **Character Filtering:** Only printable ASCII characters (codes 32–126) are allowed. Non-printable characters are reported by their numeric code in error messages.
+2. **Coordinate Tracking:** Each token records its `line` and `column` number for precise error reporting.
+3. **Implicit Concatenation:** The lexer treats characters as discrete units. The parser handles the logic of implicit concatenation.
+4. **Metacharacter Recognition:** Characters like `*`, `+`, `?`, `|`, `(`, `)`, and `.` are immediately recognised as their respective operator tokens.
+
+---
+
+## 2. Parser Details
+
+### Overview
+
+The parser consumes the token stream produced by the lexer and constructs an Abstract Syntax Tree (AST). It uses a **recursive descent** approach and strictly enforces operator precedence.
+
+### Grammar and Precedence
+
+The parser implements the following grammar (highest to lowest precedence):
+
+1. **Atom:** The basic building blocks (Literal, Dot, or a Parenthesised Expression).
+2. **Factor:** An Atom optionally followed by a quantifier (`*`, `+`, or `?`).
+3. **Term:** One or more concatenated Factors.
+4. **Expression:** One or more Terms separated by the union operator (`|`).
+
+### AST Node Structure
+
+The parser produces a tree composed of various node types (Literal, Dot, Concatenation, Union, Star, Plus, Optional).
+
+**Important Design Note:** While the base `ASTNode` class contains fields for `nullable`, `firstpos`, and `lastpos`, these are **not** populated by the parser. They are reserved for and populated by the **Stage 2 — NFA Builder**.
+
+### Error Handling
+
+The parser detects and reports the following errors with meaningful messages including line and column numbers:
+
+- **Unmatched Parentheses:** e.g., `(a|b`.
+- **Quantifier applied to nothing:** e.g., `*abc` or `(+a)`.
+- **Empty alternation branch:** Both left-side (e.g., `|abc`) and right-side (e.g., `abc|`) are explicitly detected.
+- **Unexpected Tokens:** Tokens that don't fit the expected grammar at the current position.
+
+---
+
+## Stage 2 — NFA Construction
+
+This stage converts the Abstract Syntax Tree (AST) into an ε-free Non-deterministic Finite Automaton (NFA) using Glushkov's algorithm.
+
+---
+
+### 1. Glushkov's Algorithm
+
+The algorithm produces an NFA with exactly $n+1$ states, where $n$ is the number of symbol occurrences (literals or dots) in the regex.
+
+#### Construction Steps:
+
+1. **Linearization:** Each symbol is assigned a unique integer position (1 to $n$).
+2. **Nullable:** Computes if a sub-expression can match the empty string.
+3. **Firstpos:** The set of positions that can match the first character of a string.