From Parser to Hierarchical Analysis: Comparing Natural-Language Grammar and Computer Science

本文暂无简体中文版本,正在显示英语版本。

Natural-language analysis and source-code analysis share a fundamental problem: determining structures and relations from a linear sequence of units.

In computer science, this process can be described using concepts such as Lexer (a component that performs lexical analysis), Token (a lexical unit), Parser (a component that performs syntactic analysis), and Semantic Analysis (analysis of semantic properties and constraints).

In linguistics, related concepts include 词法分析 (lexical analysis), 词类 (part of speech), 句法分析 (syntactic analysis), 层次分析 (hierarchical analysis), and 语义分析 (semantic analysis).

The two systems are not equivalent. The comparison is intended only to distinguish different levels of analysis.

Computer Science Natural Language
Source Code Linguistic Data
↓ ↓
Lexical Analysis Lexical Analysis
↓ ↓
Tokens Words / Parts of Speech
↓ ↓
Parsing Syntactic Analysis
↓ ↓
Parse Tree / AST Hierarchical Structure
↓ ↓
Semantic Analysis Semantic Analysis

The relationship between the two should be understood as a methodological similarity, not as conceptual identity.

1. From Linear Sequence to Hierarchical Structure

Source code and natural language both appear on the surface as linear sequences.

Source code:

a + b * c
我 喜欢 学习 汉语
I like study Chinese

Linear information represents order:

A → B → C → D
Linear Sequence
↓
Hierarchical Structure

This is the fundamental similarity between parsing in computer science and 层次分析 (hierarchical analysis) in grammatical analysis.

2. Lexer and Lexical Analysis: Identifying Units

In a compiler, Lexical Analysis divides a sequence of characters into Tokens.

Characters
↓
Lexical Analysis
↓
Tokens

For example:

value + 10
IDENTIFIER("value")
PLUS("+")
INTEGER(10)

Natural-language lexical analysis likewise identifies lexical units and their properties.

For example:

我喜欢学习汉语
我 / 喜欢 / 学习 / 汉语
我 Pronoun
喜欢 Verb
学习 Verb
汉语 Noun

At the methodological level:

Lexical Analysis ≈ 词法分析
Token ≈ Lexical Unit
Token Category ≈ 词类

The symbol ≈ indicates analytical similarity, not complete equivalence.

3. Tokens and Parts of Speech Do Not Determine the Complete Structure

Knowing the category of each Token is insufficient to determine the complete structure of a program.

Similarly, knowing the part of speech of each word is insufficient to determine the structure of a sentence.

A sequence such as:

Noun + Verb + Noun
Noun
├── ?
Verb
├── ?
Noun

nor does it determine the semantic relations among the three units.

In natural language, the same word may also belong to different lexical categories or perform different syntactic functions depending on its structural environment.

Classification of units is therefore a necessary part of analysis, but it is not the final result of analysis.

4. Parser and Syntactic Analysis: Determining Structure

In computer science, a Parser analyzes Tokens according to a Grammar (a formal system of structural rules) and determines how they form syntactic structures.

Tokens
↓
Parser
↓
Syntactic Structure

In linguistics, 句法分析 (syntactic analysis) determines structural relations among the units of a sentence.

Words
↓
Syntactic Analysis
↓
Syntactic Structure

At the methodological level:

Parser / Parsing ≈ 句法分析
Structure
/ \
Unit A Structure
/ \
Unit B Unit C

The important property is not the visual shape of the tree, but the structural information represented by it.

For example:

A + B + C
(A + B) + C
A + (B + C)
Parse Tree ≈ Formal Representation of Hierarchical Structure
Parsing ≈ Identification of Hierarchical Relations

Hierarchical analysis is not the AST of natural language. The two concepts only share the principle of representing structure hierarchically.

6. AST and Structural Abstraction

An Abstract Syntax Tree — AST (an abstract tree representation of syntactic structure) preserves important structural relations while omitting some formal details that are unnecessary for later processing.

For example:

a + b * c
Add
├── a
└── Multiply
├── b
└── c

One function of an AST is to transform a linear sequence into a structural model with explicit relations.

Natural-language analysis has a similar requirement. A sentence cannot be described solely by word order; the structural relations among its components must also be represented.

An AST can therefore serve as a methodological analogy for structural representation, but it should not be treated as equivalent to the syntactic structure of natural language.

7. Grammar Rules and 语法规则

In computer science, a Grammar Rule (a formal rule specifying how structures can be formed) determines which structures a Parser may construct.

For example:

Expression → Expression + Expression
Grammar Rule ≈ 语法规则
Specification
↓
Grammar
↓
Valid Structure

In the study of Natural Language:

Language Data
↓
Observation
↓
Generalization
↓
Grammatical Description

The Grammar of a programming language is usually part of its Specification. A descriptive grammar of natural language is primarily an analytical model constructed from linguistic data.

8. Pattern Is Not a Grammar Rule

A Pattern is an observable formal configuration in the data.

For example:

A + B + C
A + [B + C]
Observed Pattern
↓
Structural Hypothesis

from:

Observed Pattern
≠
Established Structure

In computer science, a Parser can normally operate according to a Grammar that has already been specified.

In natural-language research, the Grammar itself is often the model that must be constructed, described, or tested against linguistic data.

This is an important methodological difference between the two domains.

9. Candidate Structure: A Structure to Be Tested

A Pattern may produce one or more Candidate Structures (possible structural analyses requiring further evaluation).

For example:

A B C
[A B] C
A [B C]
Surface Pattern
↓
Candidate Structures
↓
Analysis

The ability to represent a Candidate Structure as a tree establishes only that a possible structural model has been constructed.

It does not establish that the model correctly represents the relations present in the linguistic data.

10. Syntactic Ambiguity and 句法歧义

Syntactic Ambiguity (the possibility of multiple syntactic analyses for the same expression) corresponds to 句法歧义 (syntactic ambiguity) in Chinese grammatical terminology.

A general model is:

Input
├── Parse A
└── Parse B

Different analyses may contain exactly the same:

  • words;
  • word order;
  • parts of speech;

while establishing different structural relations.

Therefore:

Same Tokens
≠
Same Structure

and:

Same Surface Pattern
≠
Unique Analysis

Hierarchical analysis can identify these structural possibilities. Syntactic structure, however, is not the final level of analysis.

11. Semantic Analysis and 语义分析

In a compiler, Semantic Analysis checks constraints that cannot be established from syntactic structure alone.

For example:

"hello" - 5
Subtract
├── String
└── Integer

If - is not defined for String, the expression can still be rejected.

In linguistics, 语义分析 (semantic analysis) similarly addresses information that cannot be fully established from formal structure alone.

Relevant questions include:

  • which component describes which component;
  • which participants are related to an action;
  • which entity is assigned a property;
  • whether the components are semantically compatible;
  • whether a hierarchical analysis preserves the interpretation of the expression.

At the methodological level:

Semantic Analysis ≈ Analysis of Semantic Relations and Constraints
A combines with B
What relation holds between A and B?
Words
↓
Tree

but:

Words
↓
Structure
↓
Relations

A tree is a means of representing structure. It is not independent evidence that the represented structural analysis is correct.

13. Semantic Constraints and 语义限制

In computer science, a Node may occupy a syntactically permitted position while violating a Semantic Constraint (a condition imposed on semantic validity).

For example:

Subtract
├── String
└── Integer

may have a valid syntactic shape while violating a Type requirement.

Natural-language combinations are also subject to 语义限制 (semantic constraints).

A Predicate may require the entity it describes to possess particular semantic properties. A Verb may likewise impose conditions on the participants involved in the relation it expresses.

This can be abstracted as:

Syntactic Position
↓
Semantic Requirement
↓
Compatible / Incompatible

This comparison does not imply that semantic constraints in compilers and natural languages operate through identical mechanisms.

The shared principle is:

The ability to occupy a position in a formal structure does not automatically imply semantic compatibility.

14. Why Is Pattern Insufficient to Establish an Analysis?

A Pattern provides formal evidence only.

Suppose the following sequence is observed:

A + B + C
A + [B + C]
1. Do B and C form a constituent?
2. What rule licenses this structure?
3. What syntactic relation holds between B and C?
4. What semantic relation holds between B and C?
5. Does this structure preserve the interpretation of the whole expression?

Therefore:

Pattern Match
↓
Candidate Structure
↓
Syntactic Analysis
↓
Semantic Analysis

cannot be reduced to:

Pattern Match
↓
Conclusion

A grammatical label cannot by itself establish the structure presupposed by that label. The structure must be supported by actual relations among its components.

15. Pattern + Structure + Semantics

A more complete analytical model can be represented as:

Form
↓
Pattern
↓
Candidate Structure
↓
Syntactic Relations
↓
Semantic Relations
↓
Interpretation

Each level addresses a different question.

Form

Which units are present?

Pattern

What formal distribution do these units exhibit?

Structure

How are these units organized hierarchically?

Syntactic Relations

What syntactic functions and relations do the components have?

Semantic Relations

How are the components related in meaning?

Interpretation

How is the complete structure understood in context?

A later level cannot be derived solely from the label assigned at an earlier level.

16. Conceptual Comparison

| Computer Science | Natural-Language Analysis | Point of Comparison |

|---|---|---|

| Source Code | Linguistic Data | Input data |

| Lexical Analysis | 词法分析 (lexical analysis) | Identification of units |

| Token | 词 / 词法单位 (word / lexical unit) | Unit of analysis |

| Token Category | 词类 (part of speech) | Classification of units |

| Parser | 句法分析 (syntactic analysis) | Determination of structure |

| Parse Tree | 层次结构 (hierarchical structure) | Representation of hierarchical relations |

| AST | Abstract structural model | Omission of unnecessary formal details |

| Grammar Rule | 语法规则 (grammatical rule) | Description of possible combinations |

| Syntactic Ambiguity | 句法歧义 (syntactic ambiguity) | Multiple structures for the same expression |

| Semantic Analysis | 语义分析 (semantic analysis) | Analysis of meaning relations |

| Semantic Constraint | 语义限制 (semantic constraint) | Conditions on combinations |

The pairs in this table are methodological comparisons, not definitions of equivalence.

17. Limits of the Comparison

A Programming Language is a deliberately designed formal system. A Natural Language is a system that develops through use and historical change within a linguistic community.

A compiler can usually ask:

Does this structure conform to the Specification?
What kind of system is reflected by these linguistic data?
Grammar
↓
Parser
↓
Accepted / Rejected

In natural-language research, the relationship is more complex:

Data
↓
Possible Analysis
↓
Comparison
↓
Generalization
↓
Grammatical Model

A descriptive grammar is not the “source code” of a natural language. It is a model constructed to account for linguistic data.

The value of the compiler comparison therefore lies primarily in separating analytical levels and procedures, not in treating the two systems as identical.

18. Summary: Form → Pattern → Structure → Semantics

The analysis of a linguistic expression does not end when a Pattern has been identified.

A general analytical process can be represented as:

Form
↓
Pattern
↓
Candidate Structure
↓
Syntactic Relations
↓
Semantic Relations
↓
Interpretation

Three principles follow:

Pattern ≠ Structure
Structure ≠ Semantic Validity
Possible Analysis ≠ Established Analysis

A Pattern provides a basis for proposing a structural hypothesis. A structure must be supported by syntactic relations, and those syntactic relations must in turn be evaluated against semantic relations and constraints.

This is the central methodological intersection between compiler analysis and natural-language grammatical analysis.

The next article applies this framework to a specific problem in Modern Chinese grammar: 主谓谓语句 (subject-predicate predicate sentence), including its definition, diagnostic criteria, analytical scope, and differences among competing grammatical analyses.