Why Pattern Cannot Ignore Semantics in Grammatical Analysis

本文暂无简体中文版本,正在显示英语版本。

In grammatical analysis, a Pattern can help identify structures that may be present. However, similarity in surface form is not sufficient to prove that two expressions have the same syntactic structure.

A Pattern initially provides only a Candidate Parse.

For example, given:

A + B + C
A + [B + C]
Do B and C actually form a constituent?
What syntactic relation exists between B and C?
What semantic relation exists between them?
Is C predicated of B, or of some other element?

Until these questions are resolved, [B + C] remains only a candidate structure.

This is where Semantic Analysis becomes an important part of grammatical analysis.

1. Pattern Identifies Form; It Does Not Prove Structure

Suppose we already know a structure:

A + [B + C]
X + Y + Z
X + [Y + Z]
Relation(Y, Z)
=
Relation(B, C)

Two expressions may look similar on the surface while differing in:

Constituency

Syntactic Relations

Semantic Roles

  • Semantic Relations

Therefore:

Same Surface Pattern
≠
Same Structure

A Pattern is useful for generating possible analyses. It is not, by itself, proof of structure.

2. A Compiler Perspective: Parser Acceptance Is Not Enough

The distinction between structure and semantics is particularly clear in a Compiler.

Consider the hypothetical expression:

"hello" - 5
Expression
→ Expression "-" Expression

The Parser may therefore successfully construct an Abstract Syntax Tree :

Subtract
├── String("hello")
└── Integer(5)

At the syntactic level, the structure has been identified:

Operator: -
Left: String("hello")
Right: Integer(5)

But the analysis is not finished.

Semantic Analysis and Type Checking must still determine whether these components satisfy the constraints imposed by the operator.

Suppose - is defined only for numbers:

Subtract(Number, Number) → Number
Subtract(String, Integer)
Tokens
↓
Parser
↓
AST
↓
Semantic Analysis
↓
Type Checking
↓
Valid / Invalid

Therefore:

Syntactically Constructible
≠
Semantically Valid

3. Why Is This Comparison Useful for Natural Language?

Natural language is not a programming language, nor does it operate under a formally specified Type System.

Nevertheless, both fields share a methodological problem:

Given a linear sequence of units, which elements combine with one another, and what relations exist among them?

A natural-language sentence may permit several Candidate Parses:

Input
│
├── Parse A
│
├── Parse B
│
└── Parse C

The fact that Parse A matches a familiar Pattern establishes only that:

Parse A is worth considering.

It does not establish that:

Parse A is the structure of the sentence.

The analysis must continue through:

Constituency
↓
Syntactic Relations
↓
Semantic Relations

In particular, if a proposed segmentation loses or changes an important semantic relation in the sentence, that is a reason to re-examine the analysis.

4. An Example: 我送他回国

Consider:

我送他回国。
I saw him off as he returned home.

On the surface, we can identify the sequence:

他 + 回国
他回国。
He returns home.

If we rely only on Pattern, one possible Candidate Parse is:

我 + 送 + [他回国]
Subject
├── 我
└── Predicate
├── Verb: 送
└── Object: 他回国

At the level of surface form, this analysis may initially appear plausible.

Semantic Relations, however, reveal an additional problem.

5. “Whom Does 我 送?” and “Who Returns Home?”

We can test two relations separately.

送谁?

送谁?
Whom does 我 送?
→ 他

The first relation is:

送(我, 他)
我 = Agent
他 = participant directly related to 送

Now consider the second relation.

谁回国?

谁回国?
Who returns home?
→ 他

This gives:

回国(他)
送(我, 他回国)
送(我, 他)
回国(他)

They can be visualized as:

送
├── 我
└── 他
│
└── 回国

The same element, 他, participates in two relations.

6. 兼语: One Element Participating in Two Relations

In traditional Chinese grammatical analysis, this type of construction is associated with 兼语 .

A simplified representation is:

我送他回国
│ │ │ │
│ │ │ └── Predicate
│ │ └──── Pivot
│ └─────── Verb
└───────── Subject

The crucial element is 他:

他
/ \
/ \
ObjectOf SubjectOf
↓ ↓
送 回国

A more explicit representation is:

Sentence
├── Subject: 我
└── Predicate
├── Verb: 送
├── Pivot: 他
│ ├── ObjectOf → 送
│ └── SubjectOf → 回国
└── Predicate: 回国

他 participates in a relation with 送, while simultaneously serving as the semantic subject associated with 回国.

This dual relation is what the traditional analysis of 兼语 attempts to capture.

The important point here is not that the term 兼语 must necessarily be adopted.

The methodological point is:

A syntactic analysis should preserve and explain the important relations actually present in the sentence.

7. Why Is [他回国] Not Enough to Explain the Sentence?

In:

他回国。
他 + 回国
我送他回国。
我 + 送 + [他回国]
送谁?
→ 他

This illustrates an important principle:

Possible Constituent
≠
Correct Constituent in Every Context

A sequence may form a constituent in environment A without necessarily having the same structure in environment B.

8. Semantic Relations Can Constrain a Candidate Parse

The analytical process can be represented as:

Input
↓
Pattern Recognition
↓
Candidate Parse
↓
Constituency
↓
Syntactic Relations
↓
Semantic Relations
↓
Interpretation

Suppose a Candidate Parse proposes:

Relation(A, C)
Relation(B, D)

or shows that an element inside [A + B] maintains an important direct dependency with something outside the proposed constituent.

We then have reason to reconsider:

  • the constituent boundary;
  • the type of structure;
  • the dependency relations;
  • or the analytical model itself.

Semantics does not replace Syntax.

Rather:

Semantic Relations constrain competing Candidate Parses and provide evidence for evaluating them.

9. Semantic Constraints and Type Constraints

At this point, another comparison with computer science becomes useful.

In a typed programming language:

Function : InputType → OutputType
sqrtNumber → Number
sqrt(25)
sqrt("hello")
重要(Event / Proposition) ✓
important(event / proposition)

重要 can evaluate an event or proposition.

By contrast:

好听(Auditory Entity) ✓
pleasant-to-hear(auditory entity)

好听 normally establishes a semantic relation with something that can be perceived auditorily.

A Predicate cannot combine entirely arbitrarily with every possible Semantic Argument.

If an analysis proposes:

PredicateX
Number - Number
String - Number
Semantic Mismatch
INVALID
Semantic Mismatch
↓
Context / Coercion / Metonymy
↓
Alternative Interpretation

This is one of the most important limitations of the compiler analogy.

11. Semantics Does Not Replace Syntax

If Pattern is insufficient, this does not mean that Syntax can be discarded and structure inferred entirely from meaning.

The fact that a sentence is understandable does not prove that every proposed hierarchical analysis is correct.

Given an interpretation:

Meaning M
Meaning M
→
Structure X

without structural evidence.

A complete analysis must consider:

Surface Evidence
+
Structural Evidence
+
Syntactic Relations
+
Semantic Relations

The more appropriate relationship is therefore:

Syntax ↔ Semantics
Syntax → Semantics only
Semantics → Syntax only
我送他回国。
送(我, 他)
回国(他)

A complete analysis must account for both.

If a Parse explains only:

回国(他)
送(我, 他)
送(我, 他)
Candidate Parse
↓
Preserve Syntactic Relations?
↓
Preserve Semantic Relations?
↓
Explain the Interpretation?

13. Pattern Is Still Important

Saying that Pattern is insufficient does not mean that Pattern is useless.

Patterns help us:

  • identify possible structures;
  • locate similar linguistic data;
  • generate Candidate Parses;
  • formulate grammatical rules;
  • detect unusual forms;
  • compare distributions in a Corpus.

Grammar Rules are equally important in a Compiler.

Without Grammar, a Parser would not know which structures it can construct.

But:

Grammar
↓
Parse

is not the entire compiler pipeline.

Likewise, in natural-language analysis:

Pattern
↓
Candidate Structure

is only one stage.

The methodological error occurs when:

Pattern = Proof
Pattern
=
Evidence for a Candidate Analysis

14. From Pattern Matching to Grammatical Analysis

The overall process can be summarized as:

Linguistic Data
↓
Pattern Recognition
↓
Candidate Structures
↓
Hierarchical Analysis
↓
Syntactic Relations
↓
Semantic Relations
↓
Context / Prosody
↓
Grammatical Analysis

This does not mean that natural-language analysis literally follows these steps in a rigid sequence.

It is a methodological model, not a formal algorithm for natural language.

Its purpose is to prevent an overly rapid inference:

I have seen this Pattern before
↓
Therefore I already know
the structure of this sentence

A more defensible inference is:

I have seen this Pattern before
↓
Therefore I have a Candidate Parse
↓
Now it must be tested

15. Three Distinctions That Must Be Preserved

The entire problem can be condensed into three propositions.

Pattern ≠ Structure

Pattern ≠ Structure
Structure ≠ Semantic Validity
Possible Parse ≠ Established Analysis
Source
↓
Lexer
↓
Tokens
↓
Parser
↓
AST
↓
Semantic Analysis
↓
Type Checking

Natural language does not have a rigid pipeline that corresponds exactly to this architecture.

Nevertheless, the layered mode of reasoning remains useful:

Utterance
↓
Units
↓
Patterns
↓
Candidate Structures
↓
Syntactic Relations
↓
Semantic Relations
↓
Interpretation

The common point is not:

Natural language works like a programming language.

It is:

A formal structure must still be tested against the relations it claims to represent.

A compiler makes this principle particularly visible. An AST may be syntactically constructible while its operands fail the semantic or type constraints imposed by an operator.

Natural language is much more flexible, but the methodological question remains:

This tree can be constructed.
But do the relations represented
inside this tree actually explain
the sentence?

17. Conclusion

Pattern is an important tool in grammatical analysis. It helps identify potentially relevant structures and generate Candidate Parses for further examination.

But Pattern does not prove Structure.

Structure, in turn, cannot be evaluated entirely independently of Semantic Relations.

Consider again:

我送他回国。
他 + 回国
[他回国]
送(我, 他)
回国(他)

reveals that 他 simultaneously participates in relations with both 送 and 回国.

This relationship is central to the analysis of 兼语 and demonstrates why surface Pattern cannot replace relational analysis.

A more complete analytical procedure is:

Pattern
↓
Candidate Structure
↓
Hierarchical Structure
↓
Syntactic Relations
↓
Semantic Relations
↓
Interpretation

Therefore:

Pattern ≠ Structure
Structure ≠ Semantic Validity
Possible Parse ≠ Established Analysis

The next article will apply this framework directly to 他学习很好 and 他唱歌很好听.

The question will no longer be merely whether these expressions occur or are understandable. Instead, it will examine whether 学习 and 唱歌 have the same syntactic properties, what 很好 and 很好听 are actually predicated of, and whether similar surface Patterns provide sufficient evidence for assigning the two sentences the same grammatical analysis.