Chapter 47: Word and Text Representations
Learn Machine Learning from very beginner to expert with detailed topic guidance, practical examples, practice exercises, and review questions.
What this chapter covers
This chapter contains 9 topics. Technical terms are followed by plain-language meanings in parentheses where they first appear. Code is included only when it naturally helps demonstrate the concept; architecture, workflow, governance, and comparison topics use practical scenarios instead.
47.1 One-Hot Text Representation
One-Hot Text Representation (a practical concept used within natural language processing and language-model systems). Within Chapter 47, this topic connects directly to natural language processing and language-model systems. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
Language systems convert text into tokens and numerical representations before learning or retrieving patterns. Evaluate not only fluency but also factual grounding, relevance, failure cases, context limits, and whether the model has enough trustworthy information for the task.
Example
Imagine a small machine-learning project. Use One-Hot Text Representation to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// One-Hot Text Representation
const rows = [[2,1],[4,2],[6,3],[8,4]];
const direction = [0.894, 0.447];
const projected = rows.map(row => row[0]*direction[0] + row[1]*direction[1]);
console.log(projected.map(x => x.toFixed(2)));Code explanation
- Each row begins with two numeric features.
- `direction` represents a chosen one-dimensional axis.
- The dot product projects each two-dimensional point onto that axis.
- The result shows how dimensionality reduction can compress several features into fewer numbers while preserving useful structure.
Expected result: One projected value is printed for each original row.
Practice exercise
Create a small real-world example for One-Hot Text Representation. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
47.2 Dense Embeddings
Dense Embeddings (a vector representation designed so similar items have nearby numerical representations). Within Chapter 47, this topic connects directly to natural language processing and language-model systems. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
Language systems convert text into tokens and numerical representations before learning or retrieving patterns. Evaluate not only fluency but also factual grounding, relevance, failure cases, context limits, and whether the model has enough trustworthy information for the task.
Example
Words such as 'car' and 'vehicle' can be represented by number vectors that lie closer together than unrelated words such as 'car' and 'banana'.
Coding example
// Dense Embeddings
const dot = (a,b) => a.reduce((s,x,i)=>s+x*b[i],0);
const query = [1,0.5];
const items = [[1,0],[0,1],[0.8,0.4]];
const scores = items.map(v => dot(query,v));
const best = scores.indexOf(Math.max(...scores));
console.log({ scores, bestMatch: best });Code explanation
- The query and candidate items are represented by small numeric vectors.
- A dot product produces one similarity score for each candidate.
- The largest score identifies the representation most aligned with the query.
- Modern attention and representation systems use richer versions of this same compare-and-weight idea.
Expected result: Similarity scores and the best matching item index are printed.
Practice exercise
Create a second example for Dense Embeddings. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
47.3 Word Embeddings
Word Embeddings (a vector representation designed so similar items have nearby numerical representations). Within Chapter 47, this topic connects directly to natural language processing and language-model systems. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
Language systems convert text into tokens and numerical representations before learning or retrieving patterns. Evaluate not only fluency but also factual grounding, relevance, failure cases, context limits, and whether the model has enough trustworthy information for the task.
Example
Words such as 'car' and 'vehicle' can be represented by number vectors that lie closer together than unrelated words such as 'car' and 'banana'.
Coding example
// Word Embeddings
const dot = (a,b) => a.reduce((s,x,i)=>s+x*b[i],0);
const query = [1,0.5];
const items = [[1,0],[0,1],[0.8,0.4]];
const scores = items.map(v => dot(query,v));
const best = scores.indexOf(Math.max(...scores));
console.log({ scores, bestMatch: best });Code explanation
- The query and candidate items are represented by small numeric vectors.
- A dot product produces one similarity score for each candidate.
- The largest score identifies the representation most aligned with the query.
- Modern attention and representation systems use richer versions of this same compare-and-weight idea.
Expected result: Similarity scores and the best matching item index are printed.
Practice exercise
Create a second example for Word Embeddings. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
47.4 Semantic Similarity
Semantic Similarity (a practical concept used within natural language processing and language-model systems). Within Chapter 47, this topic connects directly to natural language processing and language-model systems. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
Language systems convert text into tokens and numerical representations before learning or retrieving patterns. Evaluate not only fluency but also factual grounding, relevance, failure cases, context limits, and whether the model has enough trustworthy information for the task.
Example
Imagine a small machine-learning project. Use Semantic Similarity to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Semantic Similarity
const records = [3, 5, 7, 9, 11];
const transform = value => ({ input: value, output: value * 2 + 1 });
const results = records.map(transform);
console.log(results);Code explanation
- The sample starts with a small list of inputs so every result can be checked manually.
- `transform()` represents the main operation for this topic in a deliberately simple form.
- `map()` applies the same rule consistently to every item and returns a new result array.
- Use this pattern to focus on input, transformation, and output before replacing the toy rule with a more advanced method.
Expected result: A transformed result is printed for each input value.
Practice exercise
Create a small real-world example for Semantic Similarity. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
47.5 Contextual Embeddings
Contextual Embeddings (a vector representation designed so similar items have nearby numerical representations). Within Chapter 47, this topic connects directly to natural language processing and language-model systems. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
Language systems convert text into tokens and numerical representations before learning or retrieving patterns. Evaluate not only fluency but also factual grounding, relevance, failure cases, context limits, and whether the model has enough trustworthy information for the task.
Example
Words such as 'car' and 'vehicle' can be represented by number vectors that lie closer together than unrelated words such as 'car' and 'banana'.
Coding example
// Contextual Embeddings
const dot = (a,b) => a.reduce((s,x,i)=>s+x*b[i],0);
const query = [1,0.5];
const items = [[1,0],[0,1],[0.8,0.4]];
const scores = items.map(v => dot(query,v));
const best = scores.indexOf(Math.max(...scores));
console.log({ scores, bestMatch: best });Code explanation
- The query and candidate items are represented by small numeric vectors.
- A dot product produces one similarity score for each candidate.
- The largest score identifies the representation most aligned with the query.
- Modern attention and representation systems use richer versions of this same compare-and-weight idea.
Expected result: Similarity scores and the best matching item index are printed.
Practice exercise
Create a second example for Contextual Embeddings. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
47.6 Sentence Embeddings
Sentence Embeddings (a vector representation designed so similar items have nearby numerical representations). Within Chapter 47, this topic connects directly to natural language processing and language-model systems. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
Language systems convert text into tokens and numerical representations before learning or retrieving patterns. Evaluate not only fluency but also factual grounding, relevance, failure cases, context limits, and whether the model has enough trustworthy information for the task.
Example
Words such as 'car' and 'vehicle' can be represented by number vectors that lie closer together than unrelated words such as 'car' and 'banana'.
Coding example
// Sentence Embeddings
const dot = (a,b) => a.reduce((s,x,i)=>s+x*b[i],0);
const query = [1,0.5];
const items = [[1,0],[0,1],[0.8,0.4]];
const scores = items.map(v => dot(query,v));
const best = scores.indexOf(Math.max(...scores));
console.log({ scores, bestMatch: best });Code explanation
- The query and candidate items are represented by small numeric vectors.
- A dot product produces one similarity score for each candidate.
- The largest score identifies the representation most aligned with the query.
- Modern attention and representation systems use richer versions of this same compare-and-weight idea.
Expected result: Similarity scores and the best matching item index are printed.
Practice exercise
Create a second example for Sentence Embeddings. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
47.7 Document Embeddings
Document Embeddings (a vector representation designed so similar items have nearby numerical representations). Within Chapter 47, this topic connects directly to natural language processing and language-model systems. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
Language systems convert text into tokens and numerical representations before learning or retrieving patterns. Evaluate not only fluency but also factual grounding, relevance, failure cases, context limits, and whether the model has enough trustworthy information for the task.
Example
Words such as 'car' and 'vehicle' can be represented by number vectors that lie closer together than unrelated words such as 'car' and 'banana'.
Coding example
// Document Embeddings
const dot = (a,b) => a.reduce((s,x,i)=>s+x*b[i],0);
const query = [1,0.5];
const items = [[1,0],[0,1],[0.8,0.4]];
const scores = items.map(v => dot(query,v));
const best = scores.indexOf(Math.max(...scores));
console.log({ scores, bestMatch: best });Code explanation
- The query and candidate items are represented by small numeric vectors.
- A dot product produces one similarity score for each candidate.
- The largest score identifies the representation most aligned with the query.
- Modern attention and representation systems use richer versions of this same compare-and-weight idea.
Expected result: Similarity scores and the best matching item index are printed.
Practice exercise
Create a second example for Document Embeddings. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
47.8 Vector Databases
Vector Databases (a database optimized for storing and searching vector embeddings). Within Chapter 47, this topic connects directly to natural language processing and language-model systems. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
Language systems convert text into tokens and numerical representations before learning or retrieving patterns. Evaluate not only fluency but also factual grounding, relevance, failure cases, context limits, and whether the model has enough trustworthy information for the task.
Example
Imagine a small machine-learning project. Use Vector Databases to decide what information is needed, what step happens next, and what result should be checked.
Coding example
// Vector Databases
const a = [2, 4, 6];
const b = [1, 3, 5];
const dot = a.reduce((sum, value, i) => sum + value * b[i], 0);
const magnitude = Math.sqrt(a.reduce((sum, value) => sum + value ** 2, 0));
console.log({ dot, magnitude: magnitude.toFixed(2) });Code explanation
- The arrays `a` and `b` represent small numeric vectors so the calculation stays easy to inspect.
- `reduce()` walks through the values and combines them into one result, which is useful for many linear-algebra operations.
- The magnitude calculation squares each value, adds the squares, and takes the square root.
- The final object prints values you can compare by hand before using the same idea with larger data.
Expected result: A dot-product value and a vector magnitude are printed.
Practice exercise
Create a small real-world example for Vector Databases. Write the input, the goal, the main steps, and the result you would check. Then list one limitation or mistake a beginner should watch for.
47.9 Similarity Search
Similarity Search (a practical concept used within natural language processing and language-model systems). Within Chapter 47, this topic connects directly to natural language processing and language-model systems. The important goal is to understand what information goes into the method, what transformation or decision happens, and what output should be checked.
Language systems convert text into tokens and numerical representations before learning or retrieving patterns. Evaluate not only fluency but also factual grounding, relevance, failure cases, context limits, and whether the model has enough trustworthy information for the task.
Example
Imagine a small real-world project where Similarity Search is the main idea. Identify the input information, the decision or transformation that occurs, and the result you would inspect to decide whether the method is working correctly.
Coding example
// Similarity Search
const loss = x => (x - 10) ** 2;
const derivative = x => (loss(x + 0.0001) - loss(x - 0.0001)) / 0.0002;
let value = 0;
const rate = 0.1;
for (let step = 0; step < 6; step++) {
value -= rate * derivative(value);
}
console.log({ value: value.toFixed(3), loss: loss(value).toFixed(3) });Code explanation
- `loss()` gives a simple objective: values closer to the target produce a smaller error.
- `derivative()` estimates the slope by checking the loss just to the left and right of the current value.
- The loop repeatedly moves the value opposite the slope, which demonstrates the core idea behind gradient-based optimization.
- Printing both the final value and loss lets you confirm that the search moved toward a better solution.
Expected result: The value moves toward the target and the loss becomes smaller.
Practice exercise
Create a second example for Similarity Search. Change one important condition or input, predict how the result should change, and explain why. Then identify one limitation or common mistake a beginner should watch for.
Chapter 47 Review Questions and Answers
Q1. What is One-Hot Text Representation?
Answer: One-Hot Text Representation is a practical concept used within natural language processing and language-model systems. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q2. What is Dense Embeddings?
Answer: Dense Embeddings is a vector representation designed so similar items have nearby numerical representations. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q3. What is Word Embeddings?
Answer: Word Embeddings is a vector representation designed so similar items have nearby numerical representations. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q4. What is Semantic Similarity?
Answer: Semantic Similarity is a practical concept used within natural language processing and language-model systems. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q5. What is Contextual Embeddings?
Answer: Contextual Embeddings is a vector representation designed so similar items have nearby numerical representations. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q6. What is Sentence Embeddings?
Answer: Sentence Embeddings is a vector representation designed so similar items have nearby numerical representations. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q7. What is Document Embeddings?
Answer: Document Embeddings is a vector representation designed so similar items have nearby numerical representations. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q8. What is Vector Databases?
Answer: Vector Databases is a database optimized for storing and searching vector embeddings. In this chapter, focus on the input, the method or decision, and the result that should be checked.
Q9. What is Similarity Search?
Answer: Similarity Search is a practical concept used within natural language processing and language-model systems. In this chapter, focus on the input, the method or decision, and the result that should be checked.