LING5006

Chapter 2

From Declared Modules to Learned Organization: NLP and Its Linguistic Traditions

Lecture 02 · NLP 與語言學的分合史Lecture 02 · NLP and linguistics: a history of divergence 書籍目錄book list投影片slides

1. Introduction: When Language Became a Pipeline

Before large language models, natural language processing often looked remarkably orderly. A sentence entered a system, passed through tokenization and part-of-speech tagging, acquired a syntactic analysis, and then moved into semantic or discourse processing before reaching an application such as translation or question answering.

\[ \text{raw text} \rightarrow \text{tokens} \rightarrow \text{morphology} \rightarrow \text{syntax} \rightarrow \text{semantics} \rightarrow \text{application} \]

To a software engineer, this arrangement exemplifies sensible design: divide a difficult problem into components, define interfaces, and improve each component independently. To a linguist, however, the same sequence resembles a familiar decomposition of language into levels. Traditional NLP pipelines therefore did more than organize computation. They also encoded assumptions about what language consists of and how linguistic information should flow.

Modern large language models unsettle this picture. A Transformer does not normally receive a syntactic tree before translating a sentence. An instruction-tuned model does not visibly call a part-of-speech tagger, a semantic parser, and a discourse module before answering a question. Many tasks that once required distinct architectures now share one interface:

\[ \text{instruction} + \text{context} \rightarrow \text{model} \rightarrow \text{textual output}. \]

What, then, happened to the linguistic structure that pipelines made explicit? Did neural models show that morphemes, grammatical categories, parse trees, semantic roles, and discourse representations were unnecessary? Or did these distinctions move from declared modules into learned, distributed representations?

Answering this question requires a second caution. “Linguistics” has never named a single theory. It cannot be reduced to Chomskyan or generative grammar. Formal and generative approaches strongly influenced symbolic grammar engineering, but distributional, functional, cognitive, and usage-based traditions developed different accounts of linguistic structure. Their emphasis on frequency, context, gradient categories, constructions, and emergence from use now bears an intriguing relationship to neural language models. LLMs were not designed as implementations of those theories, yet they create a new empirical arena in which competing linguistic claims can be examined.

This chapter therefore tells two connected histories. The first traces NLP from explicit modules to latent organization. The second traces how different linguistic traditions entered, left, or re-entered computational work. The governing claim is simple:

Computational architectures do not merely process language. They embody hypotheses about which units, distinctions, and kinds of evidence are sufficient for linguistic behavior.

2. NLP as a Collection of Tasks

NLP is often introduced through tasks rather than through a unified theory of language. Many classical tasks correspond loosely to levels recognized in linguistic analysis.

Linguistic domain Typical NLP tasks Typical representation
Segmentation and orthography sentence splitting, tokenization token sequence
Morphology stemming, lemmatization, morphological analysis lemma and features
Lexical category POS tagging, word-sense disambiguation tags and senses
Syntax constituency and dependency parsing trees or graphs
Semantics semantic role labeling, semantic parsing roles, frames, logical forms
Discourse coreference and discourse parsing entity chains and relations
Pragmatics and interaction dialogue acts, intent detection communicative functions
Applications translation, QA, summarization, extraction task-specific outputs

This organization made linguistic questions operational. A linguist may ask what syntactic structure a sentence has. An NLP researcher reformulates the question as a mapping from sentence \(x\) to tree \(y\). A question about contextual grammatical category becomes:

\[ \hat{y} = \arg\max_y P(y \mid x). \]

Once researchers formulate a phenomenon as prediction over annotated examples, they can train a system, compare its predictions with a reference analysis, and quantify performance. This shift helped establish the benchmark as one of NLP’s central institutions. The Penn Treebank, for example, made large-scale syntactic annotation available for trainable and testable computational research (Marcus, Santorini, & Marcinkiewicz, 1993). Church and Mercer (1993) documented the broader empirical turn toward corpus-based computational linguistics.

The task formulation brought genuine scientific gains, but it also encouraged an ontological slide. If POS tagging, parsing, semantic role labeling, and coreference resolution become separate benchmarks, one may begin to treat the phenomena themselves as separate computational modules. A useful analytic decomposition can quietly become an architecture, and an architecture can then be mistaken for a model of the mind.

3. Linguistics Has More Than One Computational History

Histories of NLP often begin with formal grammar and then portray statistical learning as a departure from linguistics. This narrative captures part of the story but obscures the plurality of linguistic thought.

3.1 Formal and generative traditions

Formal-language theory and generative grammar encouraged researchers to represent linguistic knowledge as explicit rules and structured symbolic objects. A simple context-free rule such as

\[ S \rightarrow NP\ VP \]

asserts more than a transition between symbols. It assumes that sentences contain hierarchical constituents and that categories such as noun phrase and verb phrase support generalizations. Chomsky’s Syntactic Structures (1957) gave formal grammar a central place in linguistic theory, while later work in computational grammar developed diverse rule systems, feature structures, parsing algorithms, and interfaces to meaning.

The computer fit this outlook well. If linguistic competence could be stated as a grammar, a parser could approximate a computational implementation:

\[ \text{sentence} + \text{grammar} \rightarrow \text{structural analysis}. \]

Yet early computational linguistics was never simply an implementation arm of generative grammar. Machine translation also drew on information theory, cryptography, lexicography, engineering, and pre-generative structural analysis. Weaver’s 1949 memorandum famously framed translation partly through cryptographic and probabilistic analogies. Shannon’s theory of communication supplied probability, entropy, and sequential prediction as tools for understanding signals (Shannon, 1948). Formal grammar was a major tributary, not the whole river.

3.2 Distributional traditions

Distributional approaches begin from observable contextual patterning. Harris (1954) argued that differences in linguistic meaning correlate with differences in distribution. Firth’s often-cited account of meaning through company likewise placed contextual co-occurrence at the center of analysis (Firth, 1957). These proposals did not anticipate modern embeddings in a technically complete sense, but they supplied a durable methodological intuition: linguistic similarity can be inferred from patterns of use.

Modern distributional representations transform this intuition into geometry. If expressions occur in similar contexts, a learning system may place their vectors near one another or give them similar directions. The representation is learned from usage statistics rather than assigned through a hand-built dictionary alone.

This lineage matters because it prevents a misleading contrast between “linguistic theory” and “data.” Distributional analysis was itself a linguistic program. Corpus evidence and statistical learning did not enter NLP from a theory-free exterior.

3.3 Functional, cognitive, and usage-based traditions

Functional and cognitive approaches treat communicative function, conceptualization, and experience as central to grammar. Usage-based theories further argue that linguistic knowledge emerges through exposure to recurring forms and their functions. Frequency affects entrenchment and productivity; categories may be gradient and prototype-structured; constructions pair form with meaning at multiple levels of specificity (Bybee, 2010; Goldberg, 2006; Langacker, 1987).

These approaches differ substantially among themselves, and none maps neatly onto an LLM. Still, several properties of neural language models have strong conceptual affinities with this family of ideas:

Conceptual affinity does not establish theoretical identity. An LLM’s training corpus is not a child’s embodied and socially organized experience. Next-token prediction is not a complete theory of human language acquisition. A model may reproduce a frequency effect for reasons unlike those operating in human cognition. The scientifically useful question is therefore not whether LLMs are “generative” or “usage-based.” It is which traditions offer precise, testable explanations of what these models learn, where those explanations succeed, and where they fail.

4. Pipeline Architecture and the Layered View of Language

A characteristic late twentieth-century NLP system might take the following form:

\[ \text{Text} \rightarrow \text{Tokenizer} \rightarrow \text{Tagger} \rightarrow \text{Parser} \rightarrow \text{Semantic Analyzer} \rightarrow \text{Application}. \]

Consider the sentence Marie Curie discovered radium. Tokenization identifies a sequence. A POS tagger labels discovered as a verb. Named-entity recognition identifies Marie Curie as a person. A dependency parser may produce subject and object relations. An extraction component then produces a proposition such as:

\[ DISCOVER(\text{Marie Curie}, \text{radium}). \]

Each component creates a representation for the next. This architecture supports inspection, replacement, and reuse. A researcher can improve a parser without rebuilding every downstream application. Linguistic knowledge supplies useful inductive bias when data are limited.

However, the term modularity hides several different claims:

Kind of modularity Central question
Linguistic Are morphology, syntax, semantics, and pragmatics distinct analytic domains?
Engineering Can components be developed and replaced through stable interfaces?
Annotation Can each label set support a separate dataset and metric?
Cognitive Does the mind contain specialized or informationally encapsulated processes?

Fodor’s The Modularity of Mind (1983) concerns cognitive architecture, not software packaging. A separate POS-tagging program does not demonstrate a POS module in the brain. Conversely, evidence for some autonomy in syntactic processing does not require an NLP system to contain a parser executable. Pipelines reflect a mixture of linguistic inheritance, annotation practice, software design, and research institutions. Their visible organization cannot by itself establish the organization of human cognition.

5. The Statistical Turn: Learned Inference over Familiar Structures

During the late 1980s and 1990s, corpus-based statistical methods became increasingly dominant. This shift is often described as a transition from rules to data. The phrase is useful, but it compresses two changes that occurred at different times.

Statistical NLP initially replaced hand-written decisions inside modules while retaining many linguistic objects. A statistical tagger estimated the most probable tag sequence. A statistical parser assigned probabilities to trees. Statistical machine translation learned alignments and phrase correspondences from bilingual corpora (Brown et al., 1993). The representations remained recognizable: words, tags, constituents, dependencies, alignments, senses, and roles.

The transition can therefore be summarized as follows:

\[ \text{specified linguistic structures} + \text{hand-built inference} \]

became

\[ \text{specified linguistic structures} + \text{learned probabilistic inference}. \]

This was not the same change as end-to-end neural learning. The statistical turn weakened reliance on hand-written rules. The neural turn later weakened the requirement to declare the intermediate linguistic structures themselves.

Nor did the statistical turn simply remove linguists. Linguists and language experts remained involved in deciding annotation schemes, building corpora, analyzing errors, and defining evaluation. What changed was the locus of authority. Linguistic theory no longer determined every architectural choice, and benchmark performance gained increasing power to adjudicate among systems.

6. Why Pipelines Worked, and Why They Broke

Pipelines persisted because they solved real problems. Their outputs were inspectable. They supported a division of labor. They encoded prior knowledge that reduced the amount a learner needed to discover. They also encouraged resource reuse across applications.

Their weaknesses were equally real. The most familiar is error propagation. If a tokenizer chooses the wrong boundary, the tagger and parser receive a distorted input. If a parser commits to the wrong attachment, a later semantic component may never recover the discarded alternative.

Suppose three stages each operate at approximately 95% accuracy under highly simplified independence assumptions. The probability that all three are correct is

\[ 0.95^3 \approx 0.857. \]

The calculation is pedagogical rather than a realistic model of correlated errors, but it captures the problem: strong components do not automatically yield an equally strong pipeline.

Strict directionality also creates theoretical tension. In I saw the man with the telescope, syntax alone may not settle attachment. Semantic plausibility and world knowledge affect interpretation, as a comparison with I saw the moon with the telescope makes clear. A rigid pipeline imposes

\[ \text{syntax} \rightarrow \text{semantics}, \]

where comprehension may instead involve reciprocal constraints among syntax, meaning, discourse, and context.

End-to-end models reduce some interface errors by optimizing a shared objective. They do not abolish error. They redistribute it into representations and interactions that are harder to inspect. The engineering gain in joint optimization can therefore entail an explanatory loss in causal attribution.

7. The Neural Turn: Representation Learning

Classical machine-learning systems often depended on hand-selected features: capitalization, affixes, neighboring words, POS tags, or dictionary membership. Neural representation learning replaced a manually specified feature map

\[ \phi(x)=[f_1(x),\ldots,f_k(x)] \]

with a learned representation

\[ h=f_\theta(x). \]

The dimensions of \(h\) need not correspond individually to familiar linguistic categories. Bengio et al. (2003) demonstrated an influential neural probabilistic language model, while Collobert et al. (2011) showed how a unified neural framework could support several NLP tasks. Sequence-to-sequence learning later treated translation as a direct mapping between sequences (Sutskever, Vinyals, & Le, 2014).

The critical shift was not simply from symbols to numbers. Symbolic NLP also used numbers, and neural systems still operate over discrete inputs and outputs. The deeper change concerned who specifies the intermediate ontology. In a classical pipeline, researchers declare the categories and their interfaces. In end-to-end learning, the final objective shapes internal organization:

\[ x \rightarrow f_\theta \rightarrow y. \]

No explicit POS sequence or parse tree must intervene. This simplifies system design, but it creates a scientific mystery: where did the distinctions useful to language go?

8. Learned Structure Is Not Structure-Free

The disappearance of an explicit layer does not imply the disappearance of the information associated with that layer.

\[ \text{explicit structure} \neq \text{encoded information}. \]

A neural system may never store the symbolic relation \(nsubj(eat, child)\), yet its hidden states may support reliable recovery of subject–verb relations. This observation motivated extensive probing research. Tenney, Das, and Pavlick (2019) reported that BERT’s layers appeared to recapitulate a rough sequence resembling a classical NLP pipeline, from lower-level tagging toward higher-level semantic and discourse tasks.

Such results require care. A probe establishes that information is recoverable under a particular experimental setup. It does not automatically show that the model uses that information causally. Hewitt and Liang (2019) introduced control tasks to show that a powerful probe may learn an arbitrary mapping rather than reveal structure already present in representations. Three claims must therefore remain distinct:

\[ \text{decodable information} \neq \text{causal use} \neq \text{human-like mechanism}. \]

The pipeline may leave traces in learned representations without reappearing as a set of discrete internal modules. “Rediscovery” is an illuminating metaphor, not a literal anatomical finding.

9. The Transformer and the Collapse of Task-Specific Architecture

The Transformer provided a general architecture for modeling token relations through self-attention (Vaswani et al., 2017). Pretrained models then shifted NLP from building one model per task toward adapting one broadly trained model to many tasks. BERT showed the effectiveness of pretrained contextual representations across language-understanding benchmarks (Devlin et al., 2019). Autoregressive LLMs extended the shift by casting diverse activities as conditional generation.

Translation, entity extraction, sentiment classification, question answering, and summarization can now use nearly the same interface. This does not prove that they are cognitively identical tasks. It shows that a shared computational objective and representation can support different behaviors when instructions and context specify the required mapping.

The change affects the ontology of NLP tasks. Earlier systems often assumed that each task required its own output structure and model. Foundation models treat the task description itself as linguistic input:

\[ (\text{instruction}, \text{context}, \text{examples}) \rightarrow \text{LLM} \rightarrow \text{output}. \]

Language becomes both the object being processed and the interface that configures the processor.

10. Chinese Word Segmentation: Who Decides the Unit?

For decades in Classical Natural Language Processing (NLP), Western paradigms dominated: to process Chinese, software had to first perform Chinese Word Segmentation (CWS) (e.g., via tools like Jieba) to slice continuous character streams into Western-style “words” (cí) before feeding them into algorithms.

Chinese makes the politics and theory of units unusually visible. Written Chinese normally lacks spaces between words, and different annotation standards segment the same sequence differently. The SIGHAN Bakeoff placed standards such as CKIP, PKU, MSRA, and the Penn Chinese Treebank alongside one another, revealing that each could be internally coherent while remaining incompatible with the others (Emerson, 2005).

This variation does not make segmentation arbitrary. Boundaries can serve lexicographic, grammatical, psycholinguistic, or engineering purposes. It does show that a gold standard is an annotation convention designed for particular goals, not transparent access to a natural kind.

The debate also has a longer linguistic history. Mashi Wentong (1898) helped establish the word as a formal analytic category in Chinese grammar through engagement with European traditions. Chao (1968) emphasized that several criteria for wordhood need not converge in Chinese. Later character-based or zi-based approaches challenged the priority assigned to the word.

Neural models changed the authority structure again. Character-based models can avoid an explicit word-segmentation stage, while subword and byte-level tokenizers create units that do not correspond consistently to words, morphemes, or characters. BPE-style methods respond to corpus frequencies and an optimization trade-off involving vocabulary size and sequence length. Their units are engineering objects learned from data, although they inevitably affect linguistic behavior.

The historical progression is therefore more accurately written as:

\[ \text{linguistic analysis} \rightarrow \text{annotation standards} \rightarrow \text{learned tokenization objectives}. \]

No stage removes theory. Each relocates it. A tokenizer encodes assumptions through its training corpus, vocabulary budget, base symbols, normalization rules, and optimization procedure. The loss function does not discover the uniquely correct unit of Chinese. It selects units useful under a specified computational regime.

The irony emerged with modern deep learning and large language models (like BERT, RoBERTa, or tokenizers handling Chinese): - Engineers discovered that running models directly on single characters (字) or byte/subword representations worked significantly better than segmenting into “words.” - Explicit word segmentation caused massive out-of-vocabulary (OOV) errors, propagation of segmentation mistakes, and bloated vocabularies. - Neural networks—relying strictly on self-attention and statistical loss rather than theoretical linguistic doctrine—naturally learned higher-order semantics and syntax directly from the 字.

11. What End-to-End Learning Weakened

The move from pipelines to end-to-end learning weakened four commitments.

First, it weakened the engineering need for fixed linguistic units. Tokens may cross or subdivide word and morpheme boundaries.

Second, it weakened the requirement for explicit levels of representation. A contextual vector can jointly carry information related to lexical identity, syntax, semantic role, discourse status, and position.

Third, it weakened reliance on hand-specified symbolic categories. The model can learn distinctions that help prediction even when researchers have not named them in advance.

Fourth, it weakened unidirectional information flow. Attention-based layers repeatedly update token representations in context rather than passing a single committed analysis down an assembly line.

These changes do not amount to the rejection of generative grammar, still less of linguistics as a whole. They show that successful engineering no longer requires one particular form of explicit grammatical implementation. A linguistic theory and a task architecture answer different questions. High task performance can challenge a claim that some representation is computationally necessary for that task; it does not by itself refute the representation’s psychological reality or explanatory usefulness.

12. Usage-Based Affinities and Their Limits

Because LLMs learn from large distributions of usage, they appear at first glance to realize several usage-based commitments. Frequent expressions often receive stronger or more stable representations. Patterns can show degrees of productivity. Context shapes token representations. Construction-like form–meaning regularities can sometimes be elicited without explicit grammatical annotation.

These observations invite collaboration between NLP and usage-based linguistics. They do not license a quick theoretical victory. At least four gaps remain.

First, the training signal differs. Human learners encounter multimodal, interactive, socially situated language. Most LLM pretraining receives text and an optimization signal derived from prediction.

Second, the data regime differs. A model may see far more text than a human while receiving less grounding in perception, action, and joint attention.

Third, similar outputs can arise from different mechanisms. A frequency effect in a neural network does not establish the same cognitive explanation as a frequency effect in human processing.

Fourth, usage-based theories concern more than statistical recurrence. They make claims about categorization, analogy, entrenchment, communicative function, and the relation between grammar and experience. Each claim requires a corresponding experimental design.

The productive stance is comparative rather than celebratory. LLMs provide artificial learning systems in which exposure can be controlled, representations inspected, and counterfactual training regimes constructed. Linguistic theory supplies hypotheses more precise than general appeals to “emergence.” Together they can ask which constructions arise from which distributions, how frequency interacts with semantic coherence, and whether apparent generalization depends on memorized fragments or abstract schemas.

13. NLP and Linguistics: Separation or Reorganization?

The relationship between NLP and linguistics is often narrated as a divorce: symbolic systems depended on linguistic theory, statistical systems preferred corpora, and neural systems replaced hand-crafted structure with learned representations. There is truth in this account, but the separation is neither complete nor linear.

NLP repeatedly changes where linguistic analysis enters the scientific process. In symbolic work, linguistics often operated as specification:

\[ \text{linguistic analysis} \rightarrow \text{model architecture}. \]

In benchmark-centered statistical NLP, linguistic work often defined annotations, features, and evaluation. In current LLM research, it increasingly supplies hypotheses for analyzing a trained system:

\[ \text{training} \rightarrow \text{learned system} \rightarrow \text{linguistic explanation}. \]

This transition from linguistics as specification to linguistics as explanation is real, but incomplete. Linguistics also shapes dataset design, tokenization, multilingual evaluation, interaction protocols, and model objectives. Meanwhile, neural results feed back into debates about productivity, compositionality, gradience, typology, and the learnability of structure.

The result is less a reunion of two unified fields than a new division of explanatory labor. Formal approaches ask which structural generalizations a model captures and where it fails systematically. Distributional approaches explain how contextual statistics organize representation. Functional and usage-based approaches generate predictions about frequency, constructions, and experience. Psycholinguistics tests whether model behavior resembles human processing. Sociolinguistics and linguistic anthropology reveal which communities, registers, ideologies, and power relations training data encode.

The question is no longer whether NLP needs linguistics in the singular. It is which linguistic theories, methods, and forms of evidence help explain particular model behaviors.

14. Architectures Are Theories, but Not Complete Linguistic Theories

Every architecture contains assumptions about what information can support its objective. A parser requiring words assumes a word inventory. A dependency parser assumes labeled head–dependent relations. A Transformer assumes that repeated contextual transformations over token representations can learn useful relations. A next-token model assumes that substantial regularities can be acquired through conditional prediction:

\[ P(x_t \mid x_{

These assumptions make architectures theoretically consequential, but an architecture is not automatically a full theory of language or mind. The same architecture can implement different internal strategies after training, and the same behavior can arise from different representations. Training data and optimization shape the realized system as much as the architectural blueprint does.

The relevant contrast is therefore not linguistic theory versus no theory. It is a comparison among different sources of bias:

\[ \text{explicit linguistic bias},\quad \text{architectural bias},\quad \text{data bias},\quad \text{objective-induced bias}. \]

The task for linguistic analysis is to determine how these biases interact and which learned structures causally support behavior.

15. Conclusion: From Specification to Discovery

The history of NLP can be read as a sequence of changing answers to one question: how much linguistic organization should researchers specify before learning begins?

Symbolic NLP supplied extensive categories and rules. Statistical NLP retained many structures while learning how to select among them. Neural NLP learned features and intermediate representations. Large-scale pretraining learned broadly reusable organization before the downstream task was known. Instruction-tuned LLMs then used language itself to specify many tasks at inference time.

This sequence did not carry NLP beyond linguistics. It transformed the location of linguistic inquiry. Traditional systems displayed their commitments in modules and label inventories. LLMs require us to discover how linguistic distinctions emerge, interact, and become usable inside shared representations.

Several linguistic traditions now meet at this problem. Formal linguistics offers explicit hypotheses about hierarchy, dependency, and compositional generalization. Distributional traditions connect representation to contextual patterning. Functional, cognitive, and usage-based approaches foreground frequency, gradience, constructions, and communicative experience. None owns the LLM in advance. Each must earn explanatory force through predictions and evidence.

The central transition is thus:

\[ \boxed{\text{declared linguistic modules} \rightarrow \text{learned latent organization}} \]

and the central methodological reversal is:

\[ \boxed{\text{linguistics as specification} \rightarrow \text{linguistics as discovery and explanation}}. \]

The next chapter begins where every NLP architecture must begin: the units supplied to the model. Classical pipelines often began with words. LLMs begin with tokens. A token, however, need not be a word, morpheme, or character. Before asking what an LLM knows about language, we must therefore ask what its training procedure allows language to be made of.

Key Idea

The transition from traditional NLP to LLMs is not simply a history of rules, statistics, and neural networks. It is a shift from explicitly declared linguistic structure to learned organization. That shift did not eliminate linguistic theory. It reopened competition among formal, distributional, functional, cognitive, and usage-based explanations of what models learn.

Discussion Questions

  1. When a pipeline separates morphology, syntax, semantics, and discourse, which parts reflect linguistic analysis, engineering convenience, annotation practice, or cognitive claims?

  2. Why did statistical NLP retain more of the traditional linguistic ontology than end-to-end neural NLP?

  3. If a probe can decode syntax from a hidden state, what additional evidence would show that the model uses syntax causally?

  4. Which properties of LLMs genuinely align with usage-based theories, and which similarities may be superficial?

  5. Can downstream performance provide evidence about the psychological reality of a word or construction? Where does the inference become task-relative?

  6. What experiment could distinguish a model that memorizes recurring strings from one that learns an abstract construction?

  7. Does a unified generative interface show that translation, tagging, and question answering are the same computational task, or only that one architecture can support them?

  8. What would count as evidence that an LLM challenges a linguistic theory rather than merely one engineering implementation of it?

References

Bengio, Y., Ducharme, R., Vincent, P., & Jauvin, C. (2003). A neural probabilistic language model. Journal of Machine Learning Research, 3, 1137–1155.

Brown, P. F., Della Pietra, S. A., Della Pietra, V. J., & Mercer, R. L. (1993). The mathematics of statistical machine translation: Parameter estimation. Computational Linguistics, 19(2), 263–311.

Bybee, J. (2010). Language, Usage and Cognition. Cambridge University Press.

Chao, Y. R. (1968). A Grammar of Spoken Chinese. University of California Press.

Chomsky, N. (1957). Syntactic Structures. Mouton.

Church, K. W. (2011). A pendulum swung too far. Linguistic Issues in Language Technology, 6(5), 1–27.

Church, K. W., & Mercer, R. L. (1993). Introduction to the special issue on computational linguistics using large corpora. Computational Linguistics, 19(1), 1–24.

Collobert, R., Weston, J., Bottou, L., Karlen, M., Kavukcuoglu, K., & Kuksa, P. (2011). Natural language processing (almost) from scratch. Journal of Machine Learning Research, 12, 2493–2537.

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional Transformers for language understanding. In Proceedings of NAACL-HLT 2019 (pp. 4171–4186).

Emerson, T. (2005). The second international Chinese word segmentation bakeoff. In Proceedings of the Fourth SIGHAN Workshop on Chinese Language Processing (pp. 123–133).

Firth, J. R. (1957). A synopsis of linguistic theory, 1930–1955. In Studies in Linguistic Analysis (pp. 1–32). Blackwell.

Fodor, J. A. (1983). The Modularity of Mind. MIT Press.

Goldberg, A. E. (2006). Constructions at Work: The Nature of Generalization in Language. Oxford University Press.

Harris, Z. S. (1954). Distributional structure. Word, 10(2–3), 146–162.

Hewitt, J., & Liang, P. (2019). Designing and interpreting probes with control tasks. In Proceedings of EMNLP-IJCNLP 2019 (pp. 2733–2743).

Langacker, R. W. (1987). Foundations of Cognitive Grammar, Volume I: Theoretical Prerequisites. Stanford University Press.

Li, X., Meng, Y., Sun, X., Han, Q., Yuan, A., & Li, J. (2019). Is word segmentation necessary for deep learning of Chinese representations? In Proceedings of ACL 2019 (pp. 3242–3252).

Marcus, M. P., Santorini, B., & Marcinkiewicz, M. A. (1993). Building a large annotated corpus of English: The Penn Treebank. Computational Linguistics, 19(2), 313–330.

Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27, 379–423, 623–656.

Sutskever, I., Vinyals, O., & Le, Q. V. (2014). Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems 27.

Tenney, I., Das, D., & Pavlick, E. (2019). BERT rediscovers the classical NLP pipeline. In Proceedings of ACL 2019 (pp. 4593–4601).

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30.

Weaver, W. (1949). Translation. Memorandum, Rockefeller Foundation.