Natural-language analysis and source-code analysis share a fundamental problem: determining structures and relations from a linear sequence of units.
In computer science, this process can be described using concepts such as Lexer (a component that performs lexical analysis), Token (a lexical unit), Parser (a component that performs syntactic analysis), and Semantic Analysis (analysis of semantic properties and constraints).
In linguistics, related concepts include 词法分析 (lexical analysis), 词类 (part of speech), 句法分析 (syntactic analysis), 层次分析 (hierarchical analysis), and 语义分析 (semantic analysis).
The two systems are not equivalent. The comparison is intended only to distinguish different levels of analysis.
The relationship between the two should be understood as a methodological similarity, not as conceptual identity.
1. From Linear Sequence to Hierarchical Structure
Source code and natural language both appear on the surface as linear sequences.
Source code:
Linear information represents order:
This is the fundamental similarity between parsing in computer science and 层次分析 (hierarchical analysis) in grammatical analysis.
2. Lexer and Lexical Analysis: Identifying Units
In a compiler, Lexical Analysis divides a sequence of characters into Tokens.
For example:
Natural-language lexical analysis likewise identifies lexical units and their properties.
For example:
At the methodological level:
The symbol ≈ indicates analytical similarity, not complete equivalence.
3. Tokens and Parts of Speech Do Not Determine the Complete Structure
Knowing the category of each Token is insufficient to determine the complete structure of a program.
Similarly, knowing the part of speech of each word is insufficient to determine the structure of a sentence.
A sequence such as:
nor does it determine the semantic relations among the three units.
In natural language, the same word may also belong to different lexical categories or perform different syntactic functions depending on its structural environment.
Classification of units is therefore a necessary part of analysis, but it is not the final result of analysis.
4. Parser and Syntactic Analysis: Determining Structure
In computer science, a Parser analyzes Tokens according to a Grammar (a formal system of structural rules) and determines how they form syntactic structures.
In linguistics, 句法分析 (syntactic analysis) determines structural relations among the units of a sentence.
At the methodological level:
The important property is not the visual shape of the tree, but the structural information represented by it.
For example:
Hierarchical analysis is not the AST of natural language. The two concepts only share the principle of representing structure hierarchically.
6. AST and Structural Abstraction
An Abstract Syntax Tree — AST (an abstract tree representation of syntactic structure) preserves important structural relations while omitting some formal details that are unnecessary for later processing.
For example:
One function of an AST is to transform a linear sequence into a structural model with explicit relations.
Natural-language analysis has a similar requirement. A sentence cannot be described solely by word order; the structural relations among its components must also be represented.
An AST can therefore serve as a methodological analogy for structural representation, but it should not be treated as equivalent to the syntactic structure of natural language.
7. Grammar Rules and 语法规则
In computer science, a Grammar Rule (a formal rule specifying how structures can be formed) determines which structures a Parser may construct.
For example:
In the study of Natural Language:
The Grammar of a programming language is usually part of its Specification. A descriptive grammar of natural language is primarily an analytical model constructed from linguistic data.
8. Pattern Is Not a Grammar Rule
A Pattern is an observable formal configuration in the data.
For example:
from:
In computer science, a Parser can normally operate according to a Grammar that has already been specified.
In natural-language research, the Grammar itself is often the model that must be constructed, described, or tested against linguistic data.
This is an important methodological difference between the two domains.
9. Candidate Structure: A Structure to Be Tested
A Pattern may produce one or more Candidate Structures (possible structural analyses requiring further evaluation).
For example:
The ability to represent a Candidate Structure as a tree establishes only that a possible structural model has been constructed.
It does not establish that the model correctly represents the relations present in the linguistic data.
10. Syntactic Ambiguity and 句法歧义
Syntactic Ambiguity (the possibility of multiple syntactic analyses for the same expression) corresponds to 句法歧义 (syntactic ambiguity) in Chinese grammatical terminology.
A general model is:
Different analyses may contain exactly the same:
- words;
- word order;
- parts of speech;
while establishing different structural relations.
Therefore:
and:
Hierarchical analysis can identify these structural possibilities. Syntactic structure, however, is not the final level of analysis.
11. Semantic Analysis and 语义分析
In a compiler, Semantic Analysis checks constraints that cannot be established from syntactic structure alone.
For example:
If - is not defined for String, the expression can still be rejected.
In linguistics, 语义分析 (semantic analysis) similarly addresses information that cannot be fully established from formal structure alone.
Relevant questions include:
- which component describes which component;
- which participants are related to an action;
- which entity is assigned a property;
- whether the components are semantically compatible;
- whether a hierarchical analysis preserves the interpretation of the expression.
At the methodological level:
but:
A tree is a means of representing structure. It is not independent evidence that the represented structural analysis is correct.
13. Semantic Constraints and 语义限制
In computer science, a Node may occupy a syntactically permitted position while violating a Semantic Constraint (a condition imposed on semantic validity).
For example:
may have a valid syntactic shape while violating a Type requirement.
Natural-language combinations are also subject to 语义限制 (semantic constraints).
A Predicate may require the entity it describes to possess particular semantic properties. A Verb may likewise impose conditions on the participants involved in the relation it expresses.
This can be abstracted as:
This comparison does not imply that semantic constraints in compilers and natural languages operate through identical mechanisms.
The shared principle is:
The ability to occupy a position in a formal structure does not automatically imply semantic compatibility.
14. Why Is Pattern Insufficient to Establish an Analysis?
A Pattern provides formal evidence only.
Suppose the following sequence is observed:
Therefore:
cannot be reduced to:
A grammatical label cannot by itself establish the structure presupposed by that label. The structure must be supported by actual relations among its components.
15. Pattern + Structure + Semantics
A more complete analytical model can be represented as:
Each level addresses a different question.
Form
Which units are present?
Pattern
What formal distribution do these units exhibit?
Structure
How are these units organized hierarchically?
Syntactic Relations
What syntactic functions and relations do the components have?
Semantic Relations
How are the components related in meaning?
Interpretation
How is the complete structure understood in context?
A later level cannot be derived solely from the label assigned at an earlier level.
16. Conceptual Comparison
| Computer Science | Natural-Language Analysis | Point of Comparison |
|---|---|---|
| Source Code | Linguistic Data | Input data |
| Lexical Analysis | 词法分析 (lexical analysis) | Identification of units |
| Token | 词 / 词法单位 (word / lexical unit) | Unit of analysis |
| Token Category | 词类 (part of speech) | Classification of units |
| Parser | 句法分析 (syntactic analysis) | Determination of structure |
| Parse Tree | 层次结构 (hierarchical structure) | Representation of hierarchical relations |
| AST | Abstract structural model | Omission of unnecessary formal details |
| Grammar Rule | 语法规则 (grammatical rule) | Description of possible combinations |
| Syntactic Ambiguity | 句法歧义 (syntactic ambiguity) | Multiple structures for the same expression |
| Semantic Analysis | 语义分析 (semantic analysis) | Analysis of meaning relations |
| Semantic Constraint | 语义限制 (semantic constraint) | Conditions on combinations |
The pairs in this table are methodological comparisons, not definitions of equivalence.
17. Limits of the Comparison
A Programming Language is a deliberately designed formal system. A Natural Language is a system that develops through use and historical change within a linguistic community.
A compiler can usually ask:
In natural-language research, the relationship is more complex:
A descriptive grammar is not the “source code” of a natural language. It is a model constructed to account for linguistic data.
The value of the compiler comparison therefore lies primarily in separating analytical levels and procedures, not in treating the two systems as identical.
18. Summary: Form → Pattern → Structure → Semantics
The analysis of a linguistic expression does not end when a Pattern has been identified.
A general analytical process can be represented as:
Three principles follow:
A Pattern provides a basis for proposing a structural hypothesis. A structure must be supported by syntactic relations, and those syntactic relations must in turn be evaluated against semantic relations and constraints.
This is the central methodological intersection between compiler analysis and natural-language grammatical analysis.
The next article applies this framework to a specific problem in Modern Chinese grammar: 主谓谓语句 (subject-predicate predicate sentence), including its definition, diagnostic criteria, analytical scope, and differences among competing grammatical analyses.