The tangent space TpM of a manifold M at a point p admits three definitions: as equivalence classes of curves, as coordinate components together with a transformation law, and as derivations. All three are naturally isomorphic, and the isomorphisms do not depend on any choice of chart.
The analytic heart of the coincidence is Hadamard’s lemma. Because a smooth function decomposes as f(x)=f(a)+∑i(xi−ai)gi(x), the space of derivations is n-dimensional and ∂/∂x1∣p,…,∂/∂xn∣p form a basis.
A smooth map F:M→N induces at each point a linear map dFp:TpM→TF(p)N. In coordinates its matrix is precisely the Jacobian matrix, and the chain rule becomes the one-line identity d(G∘F)p=dGF(p)∘dFp.
Assembling all tangent spaces into TM=⨆p∈MTpM produces, in a natural way, a smooth manifold of dimension 2n. What makes this work is that the chart transitions take the form (x,v)↦(τ(x),Dτ(x)v): points move by the chart transition, vectors by its Jacobian.
What physicists describe by saying “a contravariant vector is a quantity transforming as V′μ=∂xν∂x′μVν” is exactly Definition 2, the definition by coordinate components. The four-velocity and the metric tensor of general relativity are objects living on TpM.
For a particle moving inside Rn, the velocity vector causes no trouble at all. Differentiate each component of the position γ(t)∈Rn and set γ′(t)=limh→0hγ(t+h)−γ(t). This works because Rn is a vector space, so that the difference γ(t+h)−γ(t) of two distant points means something.
Attempt the same thing for a point moving on the sphere S2 and you are stuck immediately. The “difference” of two points of the sphere does not lie on the sphere. If we regard S2 as a subset of R3, then γ(t+h)−γ(t)∈R3 does make sense, but only by borrowing the ambient space outside the sphere. A manifold need not come presented inside some large Euclidean space, and even when it is embedded we want an intrinsic definition that does not depend on the embedding.
There is a second problem, naive but troublesome. On a manifold there is no preferred coordinate system. If two charts (U,φ) and (V,ψ) cover a neighborhood of the same point p, then a velocity has two component descriptions, v=(φ∘γ)′(0) and w=(ψ∘γ)′(0). These are different tuples of numbers, yet they must describe the same physical situation. So the question “what is a tangent vector?” has to be answered in one of two ways:
define it without using coordinates, or
allow coordinates, but identify descriptions according to the transformation law that relates them.
Historically, the tradition of Riemann and of Ricci and Levi-Civita took route 2: a tensor was “a quantity carrying indices and transforming by a prescribed rule”. The modern treatment, settled in the middle of the twentieth century, takes route 1, and moreover defines a tangent vector not as “the velocity of a curve” but as “an operator that differentiates functions”. This third viewpoint looks strange at first, but its algebraic convenience is unmatched, and it leads directly into the theory of Lie groups and vector fields.
In this article we construct all three definitions and prove that they are naturally isomorphic. The goal is not to memorize one of them but to move freely among the three: curves when geometric intuition is wanted, coordinate components when a concrete computation is wanted, derivations when a proof is to be written.
Throughout, M is an n-dimensional C∞ manifold (second countable and Hausdorff). We take the definition and the handling of atlases from Smooth manifolds (the definition of a manifold(Definition 4.1)[Differentiable Manifolds]) as known. For a chart (U,φ) we write φ=(x1,…,xn) and call xi:U→R the coordinate functions. We follow the physicists’ convention of placing indices upstairs, but xi is not the i-th power of x. The partial derivative with respect to the i-th variable on Rn is written ∂i.
Let us say at once why we restrict to the C∞ category. The definition by derivations, given below, fails in the Ck setting for 1≤k<∞: the space of derivations on the algebra of germs of Ck functions turns out to be infinite-dimensional (see Remark 5.5 and the Appendix). The definitions by curves and by coordinates work for Ck with k≥1, but all three definitions agree only in the C∞ world.
A tangent vector is a local object, determined by information near p alone. To state this precisely we first set up a framework for “functions defined only near p”.
Let p∈M. On the collection of all pairs (U,f), where U is an open set containing p and f∈C∞(U), define a relation ∼ by
(U,f)∼(V,g)⟺there is an open W⊂U∩V with p∈W and f∣W=g∣W.
Then ∼ is an equivalence relation; an equivalence class is called a germ at p of a C∞ function, written [f]p or simply f. The set of all germs is denoted Cp∞(M).
Let us check that ∼ is an equivalence relation. Reflexivity and symmetry are of an obvious shape, but written out: taking W=U gives (U,f)∼(U,f), and symmetry follows because the definition is symmetric in f and g. For transitivity, if W1 witnesses (U,f)∼(V,g) and W2 witnesses (V,g)∼(V′,h), then W1∩W2 is an open set containing p on which f=g=h.
Sums, products and scalar multiples are defined on Cp∞(M) representative by representative, on the intersection of the domains (independence of the representatives follows by intersecting the sets W above), and this makes Cp∞(M) a commutative algebra over R. Moreover [f]p↦f(p) is a well-defined algebra homomorphism Cp∞(M)→R.
We begin with the most geometric definition. A curve through p is a C∞ map γ:(−ε,ε)→M with ε>0 and γ(0)=p. We define what it means for two curves to “have the same velocity at p” using a chart. Of course this is a definition only once we have checked that it does not depend on the chart.
Lemma 3.1(Agreement of velocities is chart-independent)
Let p∈M and let γ1,γ2 be C∞ curves through p. If (U,φ) and (V,ψ) are charts both containing p, then
(φ∘γ1)′(0)=(φ∘γ2)′(0)⟺(ψ∘γ1)′(0)=(ψ∘γ2)′(0).
Proof(Lemma 3.1)
Since each γj is continuous, after shrinking ε we may assume γj((−ε,ε))⊂U∩V. The map τ=ψ∘φ−1 is a C∞ diffeomorphism on φ(U∩V) (this is the compatibility condition in the atlas), and
Here Dτ(φ(p)) is the Jacobian matrix of τ, and it is invertible because τ is a diffeomorphism. Multiplying by an invertible matrix is injective, so (φ∘γ1)′(0)=(φ∘γ2)′(0) and (ψ∘γ1)′(0)=(ψ∘γ2)′(0) are equivalent.
Write Cp for the set of all C∞ curves through p. Declare γ1∼cγ2 when (φ∘γ1)′(0)=(φ∘γ2)′(0) for some chart (U,φ) containing p. By Lemma 3.1 this does not depend on the chart chosen, and it is evidently an equivalence relation. The quotient Cp/∼c is denoted TpcM, and its elements [γ] are called tangent vectors at p.
The virtue of this definition is that intuition works. Its defect is that sums and scalar multiples are not available on the spot: there is no operation on a manifold that “adds” two curves. To define the sum one has to choose a chart, add the components there, and pass back to a curve — and then check separately that the result is independent of the chart. This inconvenience is precisely the motivation for the next two definitions.
4. The second definition: coordinate components and the transformation law
We now rewrite the classical style of tensor analysis in modern language. The definition reads: a vector is a rule assigning n numbers to each coordinate system, in such a way that a change of coordinates transforms them by the Jacobian matrix.
Let Ap be the set of all pairs (φ,v) consisting of a chart (U,φ) containing p and a vector v∈Rn.
Definition 4.1(Tangent vector (definition by coordinate components))
For (φ,v),(ψ,w)∈Ap set
(φ,v)∼o(ψ,w)⟺w=D(ψ∘φ−1)(φ(p))v.
The quotient set Ap/∼o is denoted TpoM.
Proposition 4.2(The transformation law defines an equivalence relation)
∼o is an equivalence relation on Ap.
Proof(Proposition 4.2)
To shorten the notation write Jψφ=D(ψ∘φ−1)(φ(p)).
Reflexivity.φ∘φ−1 is the identity on φ(U), so Jφφ=I (the identity matrix) and v=Iv.
Symmetry. Since (ψ∘φ−1)−1=φ∘ψ−1, the formula for the derivative of an inverse map on Rn (apply the chain rule to φψ−1∘ψφ−1=id) gives Jφψ=Jψφ−1. Hence w=Jψφv implies v=Jφψw.
Transitivity. For a third chart (W,χ) we have χ∘φ−1=(χ∘ψ−1)∘(ψ∘φ−1) on φ(U∩V∩W), so the chain rule yields Jχφ=JχψJψφ. Therefore w=Jψφv and u=Jχψw imply u=Jχφv.
Note that all three properties follow from identities valid on intersections of chart domains, and those intersections are open sets containing p, hence nonempty.
∎
The set TpoM carries a natural vector space structure. Fixing one chart φ, the map v↦[(φ,v)] is a bijection Rn→TpoM (surjective by definition, injective because Jφφ=I), and we transport the linear structure of Rn through it. The transported structure is independent of the chart because the transition v↦Jψφv defining ∼o is linear. This “gluing by linear maps” is also the reason why the tangent bundle will turn out to be a vector bundle.
Remark 4.3(Correspondence with the index notation of physics)
The sentence found in textbooks of general relativity — “a contravariant vector is a quantity that transforms under a change of coordinates as V′μ=∂xν∂x′μVν” — is nothing but ∼o of Definition 4.1 written with the Einstein summation convention. Here ∂xν∂x′μ is the (μ,ν) entry of Jψφ. A covariant vector (a 1-form) is a quantity transforming by the inverse matrix, which is to say an element of the dual space of TpM.
The third definition regards a tangent vector as “an operator taking directional derivatives of functions”. Recall the directional derivative Dvf(a)=∑ivi∂if(a) on Rn: it is a real-valued linear functional f↦Dvf(a) satisfying the product rule. Conversely, we shall prove that any functional satisfying those two conditions is a directional derivative.
for all f,g∈Cp∞(M) (the Leibniz rule) is called a derivation at p. The set of all derivations is denoted TpdM.
That TpdM becomes a real vector space under (X+Y)(f)=X(f)+Y(f) and (cX)(f)=cX(f) is immediate, since both defining conditions behave linearly. Indeed, (X+Y)(fg)=X(fg)+Y(fg)=(X(f)+Y(f))g(p)+f(p)(X(g)+Y(g)).
Let X∈TpdM and, for c∈R, write c for the germ of the constant function c. Then X(c)=0.
Proof(Lemma 5.2)
Applying the Leibniz rule to 1⋅1=1 gives
X(1)=X(1)⋅1+1⋅X(1)=2X(1),
so X(1)=0. For general c, write c=c1 and use linearity: X(c)=cX(1)=0.
∎
The next result is the analytic lemma at the core of the theory. Think of it as the version of Taylor’s theorem in which the remainder is written as a smooth function.
Let A⊂Rn be an open set that is star-shaped about a∈A (that is, for every x∈A the segment {a+t(x−a):t∈[0,1]} lies in A), and let h∈C∞(A). Then there exist g1,…,gn∈C∞(A) such that for all x∈A
h(x)=h(a)+i=1∑n(xi−ai)gi(x),gi(a)=∂ih(a).
Proof(Lemma 5.3)
Fix x∈A and put u(t)=h(a+t(x−a)). Star-shapedness gives a+t(x−a)∈A for t∈[0,1], so u is defined on [0,1]; it is C∞ because h is and by the chain rule, and
gives the desired decomposition. Each gi is C∞ because the integrand is a C∞ function of (t,x)∈[0,1]×A and the interval [0,1] is compact, so differentiation under the integral sign is legitimate to every order. Finally, substituting x=a makes the integrand the constant ∂ih(a), whence gi(a)=∂ih(a).
∎
Star-shapedness is essential. When we use this lemma on a manifold we first replace the image of the chart by an open ball centered at φ(p) (shrinking the domain loses no information, since we are dealing with germs).
Let p∈M, let (U,φ) be a chart containing p with φ=(x1,…,xn), and put a=φ(p). For each i define
∂xi∂p(f):=∂i(f∘φ−1)(a)(f∈Cp∞(M)).
Then ∂/∂xi∣p∈TpdM, and (∂/∂x1∣p,…,∂/∂xn∣p) is a basis of TpdM. In particular dimTpdM=n, and every X∈TpdM is uniquely expressed as
X=i=1∑nX(xi)∂xi∂p.
Proof(Theorem 5.4)
First we check that the right-hand side does not depend on the representative of the germ. If f and f~ agree on a neighborhood W of p, then f∘φ−1 and f~∘φ−1 agree on φ(W∩U), an open neighborhood of a, so their partial derivatives at a coincide.
(1) It is a derivation. Linearity follows from linearity of partial differentiation. For the Leibniz rule, use (fg)∘φ−1=(f∘φ−1)⋅(g∘φ−1) and the product rule on Rn:
and since (f∘φ−1)(a)=f(p) and (g∘φ−1)(a)=g(p), this is the asserted identity.
(2) Linear independence. Suppose ∑ici∂/∂xi∣p=0 and apply it to the germ of the coordinate function xj∈C∞(U). Since xj∘φ−1 is the j-th coordinate function on φ(U), we have ∂i(xj∘φ−1)=δij (the Kronecker delta). Hence
0=i∑ci∂i(xj∘φ−1)(a)=i∑ciδij=cj
for every j, so c1=⋯=cn=0.
(3) They span. Take X∈TpdM and set ci=X(xi). Let f∈Cp∞(M) be arbitrary; shrinking the domain of a representative, we may assume that φ(U) is an open ball B centered at a (a ball is star-shaped). Applying Lemma 5.3 to h=f∘φ−1∈C∞(B) produces gi∈C∞(B) with
The proof of Theorem 5.4 rests entirely on Lemma 5.3, and that lemma uses smoothness of h to obtain smoothness of the gi. If one takes the same definition of derivation on the algebra of germs of Ck functions with k finite, the gi are only guaranteed to be Ck−1 and the proof collapses. The collapse is genuine: the space of derivations on germs of Ck functions is infinite-dimensional. See the Appendix for details.
Before going on, let us confirm that defining derivations on the global algebra of functions, without germs, gives the same object. This identification is used constantly in practice.
Proposition 5.6(Global derivations agree with derivations on germs)
Let TpM be the set of all R-linear maps X:C∞(M)→R satisfying the Leibniz rule X(fg)=X(f)g(p)+f(p)X(g). Then the restriction map
Φ:TpdM→TpM,Φ(Y)(f)=Y([f]p)
is a linear isomorphism.
Proof(Proposition 5.6)
On M there exist bump functions: for any open neighborhood V of p one can find χ∈C∞(M) with χ≡1 on some neighborhood of p and suppχ⊂V (this is the partition-of-unity argument in Smooth manifolds).
Linearity of Φ is clear from the definition: Φ(Y1+cY2)(f)=(Y1+cY2)([f]p)=Φ(Y1)(f)+cΦ(Y2)(f). And Φ(Y) satisfies the Leibniz rule because [fg]p=[f]p[g]p.
Injectivity. Suppose Φ(Y)=0. Given a germ [f]p, choose a representative f∈C∞(V) and, with χ as above, set f^=χf (defined to be 0 outside V). Then f^∈C∞(M) and f^=f on a neighborhood of p, so [f^]p=[f]p. Hence Y([f]p)=Y([f^]p)=Φ(Y)(f^)=0, and therefore Y=0.
Surjectivity. Let X∈TpM. We first prove locality. Suppose u,v∈C∞(M) agree on a neighborhood W of p and put w=u−v, so that w∣W=0. Choose a bump function χ with suppχ⊂W and χ≡1 near p; then χw≡0 (on W we have w=0, outside W we have χ=0). Since X(0)=0 by linearity, the Leibniz rule gives
0=X(χw)=X(χ)w(p)+χ(p)X(w)=0+X(w)
(using w(p)=0 and χ(p)=1). Hence X(u)=X(v).
Now define Y([f]p):=X(χf), where χ is a bump function supported in the domain of the representative. By locality this depends on neither the representative nor χ. Linearity is clear. For the Leibniz rule, note that χ(fg) and (χf)(χg) agree on the neighborhood where χ≡1, so
All the players are now on stage. We show that the three constructions are naturally isomorphic, “naturally” meaning that the maps giving the isomorphisms do not depend on a choice of chart.
flowchart LR
A["Equivalence classes of curves T_p^c M<br/>curves carrying a velocity"] -->|"Λ: send a curve to its components (φ∘γ)'(0)"| B["Coordinate components T_p^o M<br/>n numbers transforming by the Jacobian"]
B -->|"Θ: send components v to Σ vⁱ ∂/∂xⁱ"| C["Derivations T_p^d M<br/>functionals satisfying the Leibniz rule"]
A -->|"Ξ: send a curve to f ↦ (f∘γ)'(0)"| C
The three definitions and the natural maps linking them; the triangle commutes.
Theorem 6.1(Equivalence of the three tangent spaces)
Let p∈M. Each of the following maps is a bijection, and Ξ=Θ∘Λ.
Λ:TpcM→TpoM, Λ([γ])=[(φ,(φ∘γ)′(0))], where (U,φ) is any chart containing p.
Θ:TpoM→TpdM, Θ([(φ,v)])=∑i=1nvi∂xi∂p.
Ξ:TpcM→TpdM, Ξ([γ])(f)=(f∘γ)′(0).
Moreover Θ is a linear isomorphism, and if TpcM is given the linear structure transported by Ξ, then all three are isomorphic as vector spaces. From now on we identify them and write TpM.
Proof(Theorem 6.1)
Λ is well defined and bijective. If γ1∼cγ2, then by definition their velocities agree in the same chart, so the value of Λ is independent of the representative. Independence of the chart follows from the identity (ψ∘γ)′(0)=Jψφ(φ∘γ)′(0) established in the proof of Lemma 3.1: it says exactly that (φ,(φ∘γ)′(0))∼o(ψ,(ψ∘γ)′(0)). Injectivity: reading Λ([γ1])=Λ([γ2]) through the representative in the chart φ gives (φ∘γ1)′(0)=(φ∘γ2)′(0), i.e. γ1∼cγ2. Surjectivity: given [(φ,v)], put a=φ(p) and define
γ(t)=φ−1(a+tv)(∣t∣<ε).
Since φ(U) is open, ε>0 can be taken small enough that a+tv∈φ(U); then γ is C∞ as a composition of C∞ maps, γ(0)=p, and (φ∘γ)(t)=a+tv gives (φ∘γ)′(0)=v.
Θ is well defined. Suppose (φ,v)∼o(ψ,w), i.e. w=Jψφv. Write ψ=(y1,…,yn), τ=ψ∘φ−1 and a=φ(p), so that the (j,i) entry of Jψφ is ∂iτj(a). For any germ f, applying the chain rule on Rn to f∘φ−1=(f∘ψ−1)∘τ gives
so the value of Θ is independent of the representative.
Θ is a linear isomorphism. Fixing a chart φ, the correspondence TpoM∋[(φ,v)]↔v∈Rn is a linear isomorphism (this is exactly how the linear structure was defined in §4), and under it Θ becomes v↦∑ivi∂/∂xi∣p. By Theorem 5.4 the family {∂/∂xi∣p} is a basis of TpdM, so this map is a linear isomorphism.
Ξ=Θ∘Λ. Let γ be a curve through p and v=(φ∘γ)′(0). For any germ f we have f∘γ=(f∘φ−1)∘(φ∘γ), so the chain rule gives
The right-hand side is Θ(Λ([γ]))(f). Incidentally this also shows that Ξ([γ]) is a derivation, being in the image of Θ∘Λ. That Ξ is a bijection follows, as a composition, from the bijectivity of Λ and Θ.
Equip M=Rn with the atlas whose only chart is the identity map id. The coordinate functions are the usual xi, and ∂/∂xi∣a is the ordinary partial derivative f↦∂if(a). By Theorem 5.4, TaRn is n-dimensional, and the map sending v∈Rn to ∑ivi∂/∂xi∣a is a linear isomorphism Rn∼TaRn. This isomorphism is canonical — no chart has to be chosen — so from now on we identify TaRn with Rn. On the side of curves it is the familiar correspondence [γ]↦γ′(0).
Example 6.3(Rewriting the basis in polar coordinates)
Take M=R2 and U={(x,y):x>0}, and put the polar chart ψ=(r,θ) on U by prescribing its inverse
ψ−1(r,θ)=(rcosθ,rsinθ),r>0,−2π<θ<2π.
We use the transformation formula from the proof of Theorem 6.1, with polar coordinates as the source and Cartesian coordinates as the target. The Jacobian matrix of τ=(Cartesian)∘(polar)−1 is
As a check, take f(x,y)=x2+y2. In polar coordinates f∘ψ−1(r,θ)=r2, so ∂/∂r∣p(f)=2r=22 and ∂/∂θ∣p(f)=0. Computing instead with the right-hand sides,
21(2x+2y)(1,1)=24=22,(−2x+2y)(1,1)=0,
in agreement. Note that ∂/∂θ∣p is not a “unit vector”. We have not yet introduced any notion of length, but its Cartesian components are (−1,1), of Euclidean length 2=r. That a coordinate basis need not be orthonormal is the first thing to get used to when working with curvilinear coordinates.
The greatest dividend of introducing tangent spaces is that a smooth map can be linearized at each point. With the definition by derivations, the differential can be written down with astonishing brevity.
Let F:M→N be a C∞ map, p∈M and q=F(p). For X∈TpM define dFp(X) by
(dFp(X))(g):=X(g∘F)(g∈Cq∞(N)).
The map dFp:TpM→TqN is called the differential of F at p (also the pushforward, written F∗p).
Let us first check that the definition makes sense. If g is a C∞ function on a neighborhood V of q, then F−1(V) is an open neighborhood of p by continuity of F, and g∘F∈C∞(F−1(V)), so [g∘F]p∈Cp∞(M) is well defined. Replacing the representative of g on a neighborhood of q does not change g∘F on a neighborhood of p, so we obtain a map of germs F∗:Cq∞(N)→Cp∞(M), [g]↦[g∘F], and it is an algebra homomorphism.
In Definition 7.1, dFp(X) is a derivation at q, and dFp:TpM→TqN is linear.
Proof(Proposition 7.2)
Linearity of dFp(X) holds because F∗ is linear and X is linear. For the Leibniz rule, use that F∗ preserves products, (gh)∘F=(g∘F)(h∘F), together with the Leibniz rule for X:
In the situation above, if X=Ξ([γ])∈TpM then dFp(X)=Ξ([F∘γ]). In the language of curves, then, the differential is nothing but the operation of pushing a curve forward by F.
Proof(Proposition 7.3)
Being a composition of C∞ maps, F∘γ is a C∞ curve through q. For any g∈Cq∞(N),
As h is arbitrary, the identity follows. For the identity map, d(id)p(X)(f)=X(f∘id)=X(f).
∎
Corollary 7.5(The differential of a diffeomorphism is an isomorphism)
If F:M→N is a diffeomorphism, then dFp:TpM→TF(p)N is a linear isomorphism at every p, with (dFp)−1=d(F−1)F(p). In particular, corresponding points of diffeomorphic manifolds have tangent spaces of the same dimension.
Proof(Corollary 7.5)
Applying Theorem 7.4 to F−1∘F=idM and F∘F−1=idN gives
Proposition 7.6(The matrix of the differential is the Jacobian)
Let F:M→N be C∞, let (U,φ) be a chart of M containing p (coordinates xi, dimM=n) and (V,ψ) a chart of N containing q=F(p) (coordinates yj, dimN=m), and put F=ψ∘F∘φ−1 and a=φ(p). Then
dFp(∂xi∂p)=j=1∑m∂xi∂Fj(a)∂yj∂q.
That is, the matrix of dFp with respect to the bases {∂/∂xi∣p} and {∂/∂yj∣q} is the Jacobian matrix DF(a) of the local representative F.
(using F(a)=ψ(q)). Since g is arbitrary, the claim follows.
∎
This proposition guarantees that the abstract Definition 7.1 is the same thing as the classical “linear approximation by the Jacobian matrix”. Writing Theorem 7.4 out through Proposition 7.6 returns the chain rule as a product of matrices, D(G∘F)(a)=DG(F(a))DF(a).
Let Sn={u∈Rn+1:∥u∥=1} be an embedded submanifold of Rn+1 and ι:Sn↪Rn+1 the inclusion. Under the identification TpRn+1≅Rn+1 of Example 6.2,
dιp(TpSn)=p⊥={v∈Rn+1:⟨p,v⟩=0}.
(One inclusion.) Given [γ]∈TpSn, the curve ι∘γ lies in Rn+1 and satisfies ∥γ(t)∥2=1. Differentiating in t (using bilinearity of the inner product and the product rule),
(The other inclusion.) Let v∈p⊥ with v=0, put e=v/∥v∥ and ω=∥v∥, and define
γ(t)=(cosωt)p+(sinωt)e.
Since p and e are orthonormal, ∥γ(t)∥2=cos2ωt+sin2ωt=1, so γ is a curve in Sn; as Sn is an embedded submanifold, it is C∞ also as a map into Sn. From γ(0)=p and γ′(0)=ωe=v we conclude that v lies in the image. The vector v=0 is obtained as 0=dιp(0).
(Conclusion.) The map dιp is linear (Proposition 7.2), and its image both contains and is contained in p⊥, hence equals p⊥. Since dimTpSn=n=dimp⊥ (Theorem 5.4), a surjective linear map between spaces of equal dimension is also injective, so dιp is injective and TpSn≅p⊥. The familiar picture of “a plane touching the sphere” is this isomorphism drawn inside Rn+1 after translating it so that its base point is p.
Example 7.8(The tangent space of the orthogonal group at the identity)
the space of skew-symmetric matrices. Indeed, a curve γ in O(n) with γ(0)=I satisfies γ(t)Tγ(t)=I, and differentiating at t=0 gives
γ′(0)Tγ(0)+γ(0)Tγ′(0)=γ′(0)T+γ′(0)=0.
Conversely, if AT=−A then γ(t)=exp(tA) takes values in O(n), because γ(t)T=exp(tAT)=exp(−tA)=γ(t)−1, and γ(0)=I, γ′(0)=A. The space of skew-symmetric matrices has dimension n(n−1)/2, matching dimO(n), so the same dimension count as in Example 7.7 settles the equality.
Now that we have a tangent space at each point, we bundle them together over all points. A vector field on a manifold is “a rule assigning to each point a tangent vector at that point”, but to discuss its smoothness the set of all tangent vectors must itself be a manifold.
Let M be an n-dimensional C∞ manifold. As a set, put
TM=p∈M⨆TpM={(p,X):p∈M,X∈TpM}
and call π:TM→M, π(p,X)=p, the projection. We call TM the tangent bundle of M, and π−1(p)=TpM the fiber over p.
Schematic picture of the tangent bundle: over each point of the base M stands a fiber T_pM, and a vector field is drawn as a section threading through them.
Theorem 8.2(The tangent bundle is a 2n-dimensional manifold)
Let M be an n-dimensional C∞ manifold. For a C∞ atlas {(Uα,φα)} of M, writing φα=(xα1,…,xαn), define
Then there is exactly one topology and C∞ structure on TM having {(π−1(Uα),φα)} as an atlas, and TM becomes a C∞ manifold of dimension 2n. Moreover π:TM→M is a surjective C∞ submersion.
Proof(Theorem 8.2)
(1) Each φα is a bijection. By Theorem 5.4, every tangent vector at p∈Uα has a unique component expression in the basis {∂/∂xαi∣p}, so π−1(Uα)→φα(Uα)×Rn is bijective.
(2) Transition maps. Let Uαβ=Uα∩Uβ=∅ and put τ=φβ∘φα−1. The set φα(π−1(Uαβ))=φα(Uαβ)×Rn is open in R2n. By the change-of-basis formula established in the proof of Theorem 6.1,
φβ∘φα−1(x,v)=(τ(x),Dτ(x)v).
Since τ is C∞, every entry of x↦Dτ(x) is C∞, and Dτ(x)v is C∞ in x and linear in v, hence C∞ altogether. The inverse map has the same form with α and β interchanged, so the transition maps are diffeomorphisms.
(3) Topology. We check that B={φα−1(W):α,W⊂φα(Uα)×Rn open} is a basis for a topology on TM. It covers TM. Take a point ξ in the intersection of two members φα−1(W1) and φβ−1(W2); then ξ∈π−1(Uαβ), and by (2) the set φα(φα−1(W1)∩φβ−1(W2))=W1∩(φβ∘φα−1)−1(W2) is open; calling it W3 we get ξ∈φα−1(W3)⊂ the intersection. So B is a basis, and in the topology it generates each φα is a homeomorphism.
(4) Hausdorff. Let ξ=η in TM. If π(ξ)=π(η)=p, choose α with p∈Uα; then φα(ξ)=φα(η), and these can be separated inside φα(Uα)×Rn because R2n is Hausdorff; the preimages are the required open sets. If π(ξ)=π(η), then since M is Hausdorff there are disjoint open sets U∋π(ξ) and V∋π(η), and π−1(U) and π−1(V) are disjoint open sets (π is continuous in the topology of (3) because its local representative is the continuous map (x,v)↦x).
(5) Second countability. As M is second countable, there is a countable subatlas {(Uαk,φαk)}k∈N. Each π−1(Uαk) is homeomorphic to an open subset of R2n and hence second countable, and a space covered by countably many second countable open subspaces is second countable.
(6) The C∞ structure and π. By (2), {(π−1(Uα),φα)} is a C∞ atlas, and the maximal atlas containing it determines the C∞ structure. Both the topology and the atlas are forced by the requirement that each φα be a diffeomorphism, so they are unique. The local representative of π is the projection φα∘π∘φα−1(x,v)=x, which is C∞ with surjective differential (its Jacobian matrix is (In0)), so π is a submersion. Surjectivity follows because every TpM contains the zero vector.
Let S1⊂R2 and identify TpS1≅p⊥ using Example 7.7, where p=(p1,p2). The line p⊥ is one-dimensional, with basis Jp:=(−p2,p1) (indeed ⟨p,Jp⟩=−p1p2+p2p1=0 and ∥Jp∥=1=0). So define
Ψ:S1×R→TS1,Ψ(p,t)=(p,tJp).
Then Ψ is a bijection, since t↦tJp is an isomorphism on each fiber. For smoothness, take the chart θ↦(cosθ,sinθ) of S1: the local representative of Ψ becomes (θ,t)↦(θ,t). Indeed, in these coordinates ∂/∂θ∣p=(−sinθ,cosθ)=Jp, so the component of tJp is exactly t. The inverse is C∞ for the same reason, so Ψ is a diffeomorphism and TS1≅S1×R. Equivalently, S1 carries a nowhere vanishing vector field, namely p↦Jp.
What happens in Example 8.4 is not typical of all manifolds. A manifold with TM≅M×Rn is called parallelizable, and S2 is not parallelizable. This is a consequence of the hairy ball theorem (Poincaré–Brouwer): every continuous vector field on S2 vanishes somewhere. Among spheres, only S1, S3 and S7 are known to be parallelizable (Bott–Milnor and Kervaire, 1958). That the question of triviality of the tangent bundle contains deep topology is one of the reasons for introducing TM at all.
On M=R2, besides the usual coordinates (x,y), introduce the global chart ψ=(u,v) given by u=x+y and v=x−y. At a point p, express ∂/∂u∣p and ∂/∂v∣p in terms of ∂/∂x∣p and ∂/∂y∣p. Then check the expression on the function f(x,y)=xy.
Solution
Since ψ−1(u,v)=(2u+v,2u−v), the Jacobian matrix of τ=φ∘ψ−1 (with φ the Cartesian chart) is
Dτ=(1/21/21/2−1/2).
By the transformation formula in the proof of Theorem 6.1 (namely ∂/∂ui=∑j∂iτj⋅∂/∂xj),
Now the check. In the new coordinates f(x,y)=xy becomes f∘ψ−1(u,v)=2u+v⋅2u−v=4u2−v2, so ∂/∂u∣p(f)=u/2 and ∂/∂v∣p(f)=−v/2. Computing with the right-hand sides gives 21(y+x)=2x+y=2u and 21(y−x)=−2x−y=−2v, in agreement.
Let X∈TpM be a derivation and let f,g∈Cp∞(M) satisfy f(p)=g(p)=0. Show that X(fg)=0. Then, setting mp={f∈Cp∞(M):f(p)=0} and mp2={∑k=1Nfkgk:fk,gk∈mp}, show that X↦X∣mp gives a linear isomorphism TpM→(mp/mp2)∗.
Solution
The first claim is an immediate consequence of the Leibniz rule: X(fg)=X(f)g(p)+f(p)X(g)=X(f)⋅0+0⋅X(g)=0. By linearity, X vanishes on every element of mp2, these being finite sums.
Now the isomorphism. Restricting X∈TpM to mp gives, by the above, a map vanishing on mp2, hence an induced linear functional Xˉ on the quotient mp/mp2. The correspondence X↦Xˉ is clearly linear, restriction and passage to a quotient both being linear operations.
Injectivity. Suppose Xˉ=0, i.e. X∣mp=0. Any f∈Cp∞(M) decomposes as f=f(p)+(f−f(p)), with the second term in mp. By Lemma 5.2 we have X(f(p))=0, so X(f)=X(f−f(p))=0. Hence X=0.
Surjectivity. Given λ∈(mp/mp2)∗, define X(f):=λ([f−f(p)]). Linearity is clear. For the Leibniz rule, decompose
fg−f(p)g(p)=f(p)(g−g(p))+g(p)(f−f(p))+(f−f(p))(g−g(p))
and note that the last term lies in mp2, hence is killed by λ. That this X induces λ is clear from the definition.
Combined with Theorem 5.4, this shows dimmp/mp2=n, and mp/mp2 is identified with the cotangent space Tp∗M. From this angle, Lemma 5.3 says that every element of mp is a linear combination of the coordinate functions plus an element of mp2.
The set GL(n,R)={A∈Matn(R):detA=0} is open in Matn(R)≅Rn2 and hence a manifold of dimension n2. Under the identification TIGL(n,R)≅Matn(R), show that the differential of det:GL(n,R)→R at the identity is d(det)I(A)=trA.
Solution
Since GL(n,R) is open in Rn2, taking the inclusion as a chart gives TIGL(n,R)≅Matn(R) (the same argument as in Example 6.2). For A∈Matn(R) take the curve γ(t)=I+tA. As det is continuous, detγ(t)=0 for ∣t∣ small enough, so γ is a curve in GL(n,R) with γ(0)=I and γ′(0)=A. By Proposition 7.3, d(det)I(A)=(det∘γ)′(0).
Collect the terms of degree at most 1 in t. If σ is not the identity permutation, then σ(i)=i for at least two values of i (a permutation moves at least two points), so the corresponding product is divisible by t2. The identity permutation contributes ∏i(1+tAii)=1+t∑iAii+O(t2). Hence
det(I+tA)=1+ttrA+O(t2),
and differentiating in t at t=0 gives d(det)I(A)=trA.
Let M be a connected C∞ manifold, N a C∞ manifold and F:M→N a C∞ map. Show that if dFp=0 for every p∈M, then F is constant.
Solution
Fix p0∈M arbitrarily, put q0=F(p0), and show that S=F−1(q0) is nonempty, open and closed. Connectedness of M (characterization of connectedness(Proposition 3.1)[連結性]: a nonempty subset that is both open and closed must be the whole space) then gives S=M, i.e. F is constant.
S=∅ because p0∈S. And S is closed because {q0} is closed (N being Hausdorff) and F is continuous.
Now we show S is open. Take p∈S and choose a chart (U,φ) containing p and a chart (V,ψ) containing F(p)=q0 such that F(U)⊂V and φ(U) is an open ball, hence connected (by continuity of F it suffices to shrink U to U∩F−1(V)). Put F=ψ∘F∘φ−1:φ(U)→ψ(V). By Proposition 7.6, the condition dFx=0 for x∈U is equivalent to the vanishing of the Jacobian matrix DF(φ(x)). By hypothesis this holds on all of φ(U), so all partial derivatives of every component of F vanish. Since φ(U) is a connected open set, F is constant (componentwise, either join points by segments and apply the mean value theorem, or quote it as a corollary in Differentiation in several variables). As F(φ(p))=ψ(q0), we get F≡ψ(q0), that is F∣U≡q0, so U⊂S. Hence S is open.
Dropping the hypothesis that M is connected destroys the conclusion. For instance take M=R⊔R (the disjoint union of two lines) and N=R, and let F be 0 on one copy and 1 on the other: then dFp=0 at every point, yet F is not constant.
Matsumoto Yukio, Tayōtai no Kiso, University of Tokyo Press, 1988 (in Japanese) — an elementary introduction to tangent vectors and tangent spaces. A representative example of the exposition that starts from the definition by curves.
J. M. Lee, Introduction to Smooth Manifolds, 2nd ed., Springer GTM 218, 2013 — Chapter 3 (Tangent Vectors). The definition by derivations, the differential, and the construction of the tangent bundle are developed in almost the same order as in this article.
F. W. Warner, Foundations of Differentiable Manifolds and Lie Groups, Springer GTM 94, 1983 — Chapter 1. A concise account of the definition by derivations and of why smoothness is needed.
M. Spivak, A Comprehensive Introduction to Differential Geometry, Vol. 1, 3rd ed., Publish or Perish, 1999 — Chapter 3. A detailed side-by-side comparison of the three definitions, with the role of Hadamard’s lemma made explicit.
R. Bott and J. Milnor, “On the parallelizability of the spheres”, Bulletin of the American Mathematical Society 64 (1958), 87–89 — the result on parallelizability of spheres mentioned in Remark 8.5.
Appendix: Why the space of derivations grows in the Ck setting
Where the problem lies. What was essential in the proof of Theorem 5.4 was Lemma 5.3, that is, the fact that an element of mp falls into mp2 once linear combinations of the coordinate functions are removed. As we saw in Exercise 9.2, the space of derivations at p is identified with (mp/mp2)∗. So the statement “the space of derivations is n-dimensional” and the statement ”mp/mp2 is n-dimensional” are one and the same.
What goes wrong for C1. Take n=1 and p=0, and consider the algebra C01(R) of germs of C1 functions. Let m be the germs vanishing at 0 and m2 the finite sums of products of such. For 0<α<1 consider fα(x)=∣x∣1+α. It is C1 (its derivative fα′(x)=(1+α)sgn(x)∣x∣α is continuous at 0 with value 0) and lies in m. On the other hand, if g,h∈m are C1, then the mean value theorem gives ∣g(x)∣≤C∣x∣ and ∣h(x)∣≤C∣x∣ near 0, so any finite sum ∑kgkhk is O(x2). But ∣x∣1+α/x2→∞ as x→0 when α<1, so fα∈/m2. Moreover, the same comparison of growth rates shows that the fα for distinct α are linearly independent in m/m2. Hence m/m2 is infinite-dimensional, and so is its dual, the space of derivations.
Conclusion. In the C∞ world a small miracle occurs: mp/mp2 collapses to exactly n dimensions, and what guarantees this is Lemma 5.3. When working with Ck manifolds for finite k, the standard prescription is to adopt Definition 3.2 or Definition 4.1 as the definition of the tangent space and to give up the characterization by derivations. A detailed discussion of this state of affairs can be found in Chapter 1 of Warner.
Related articles
ベクトル場と微分形式:切断・外積代数・外微分ベクトル場を接バンドルの切断として定義し、余接バンドルから k 次微分形式と外積代数を構成する。外微分 d の存在と一意性を公理から証明し、d∘d=0 が grad・rot・div の恒等式を統一することを示す。