Skip to content

Tangent Spaces and the Tangent Bundle: Three Ways to Define Velocity on a Curved Space

Prerequisite:Differentiable Manifolds: Charts and Atlases That Bring Calculus to Curved Spaces

Raw
  • The tangent space TpMT_pM of a manifold MM at a point pp admits three definitions: as equivalence classes of curves, as coordinate components together with a transformation law, and as derivations. All three are naturally isomorphic, and the isomorphisms do not depend on any choice of chart.
  • The analytic heart of the coincidence is Hadamard’s lemma. Because a smooth function decomposes as f(x)=f(a)+i(xiai)gi(x)f(x) = f(a) + \sum_i (x^i - a^i) g_i(x), the space of derivations is nn-dimensional and /x1p,,/xnp\partial/\partial x^1|_p, \ldots, \partial/\partial x^n|_p form a basis.
  • A smooth map F:MNF : M \to N induces at each point a linear map dFp:TpMTF(p)NdF_p : T_pM \to T_{F(p)}N. In coordinates its matrix is precisely the Jacobian matrix, and the chain rule becomes the one-line identity d(GF)p=dGF(p)dFpd(G \circ F)_p = dG_{F(p)} \circ dF_p.
  • Assembling all tangent spaces into TM=pMTpMTM = \bigsqcup_{p \in M} T_pM produces, in a natural way, a smooth manifold of dimension 2n2n. What makes this work is that the chart transitions take the form (x,v)(τ(x),Dτ(x)v)(x, v) \mapsto (\tau(x), D\tau(x)v): points move by the chart transition, vectors by its Jacobian.
  • What physicists describe by saying “a contravariant vector is a quantity transforming as Vμ=xμxνVνV'^{\mu} = \frac{\partial x'^{\mu}}{\partial x^{\nu}} V^{\nu}” is exactly Definition 2, the definition by coordinate components. The four-velocity and the metric tensor of general relativity are objects living on TpMT_pM.

1. Motivation: putting arrows on a curved space

Section titled “1. Motivation: putting arrows on a curved space”

For a particle moving inside Rn\mathbb{R}^n, the velocity vector causes no trouble at all. Differentiate each component of the position γ(t)Rn\gamma(t) \in \mathbb{R}^n and set γ(t)=limh0γ(t+h)γ(t)h\gamma'(t) = \lim_{h \to 0} \frac{\gamma(t+h) - \gamma(t)}{h}. This works because Rn\mathbb{R}^n is a vector space, so that the difference γ(t+h)γ(t)\gamma(t+h) - \gamma(t) of two distant points means something.

Attempt the same thing for a point moving on the sphere S2S^2 and you are stuck immediately. The “difference” of two points of the sphere does not lie on the sphere. If we regard S2S^2 as a subset of R3\mathbb{R}^3, then γ(t+h)γ(t)R3\gamma(t+h) - \gamma(t) \in \mathbb{R}^3 does make sense, but only by borrowing the ambient space outside the sphere. A manifold need not come presented inside some large Euclidean space, and even when it is embedded we want an intrinsic definition that does not depend on the embedding.

There is a second problem, naive but troublesome. On a manifold there is no preferred coordinate system. If two charts (U,φ)(U, \varphi) and (V,ψ)(V, \psi) cover a neighborhood of the same point pp, then a velocity has two component descriptions, v=(φγ)(0)v = (\varphi \circ \gamma)'(0) and w=(ψγ)(0)w = (\psi \circ \gamma)'(0). These are different tuples of numbers, yet they must describe the same physical situation. So the question “what is a tangent vector?” has to be answered in one of two ways:

  1. define it without using coordinates, or
  2. allow coordinates, but identify descriptions according to the transformation law that relates them.

Historically, the tradition of Riemann and of Ricci and Levi-Civita took route 2: a tensor was “a quantity carrying indices and transforming by a prescribed rule”. The modern treatment, settled in the middle of the twentieth century, takes route 1, and moreover defines a tangent vector not as “the velocity of a curve” but as “an operator that differentiates functions”. This third viewpoint looks strange at first, but its algebraic convenience is unmatched, and it leads directly into the theory of Lie groups and vector fields.

In this article we construct all three definitions and prove that they are naturally isomorphic. The goal is not to memorize one of them but to move freely among the three: curves when geometric intuition is wanted, coordinate components when a concrete computation is wanted, derivations when a proof is to be written.

2. Preliminaries: notation and germs of functions

Section titled “2. Preliminaries: notation and germs of functions”

Throughout, MM is an nn-dimensional CC^{\infty} manifold (second countable and Hausdorff). We take the definition and the handling of atlases from Smooth manifolds (the definition of a manifold(Definition 4.1)[Differentiable Manifolds]) as known. For a chart (U,φ)(U, \varphi) we write φ=(x1,,xn)\varphi = (x^1, \ldots, x^n) and call xi:URx^i : U \to \mathbb{R} the coordinate functions. We follow the physicists’ convention of placing indices upstairs, but xix^i is not the ii-th power of xx. The partial derivative with respect to the ii-th variable on Rn\mathbb{R}^n is written i\partial_i.

Let us say at once why we restrict to the CC^{\infty} category. The definition by derivations, given below, fails in the CkC^k setting for 1k<1 \le k < \infty: the space of derivations on the algebra of germs of CkC^k functions turns out to be infinite-dimensional (see Remark 5.5 and the Appendix). The definitions by curves and by coordinates work for CkC^k with k1k \ge 1, but all three definitions agree only in the CC^{\infty} world.

A tangent vector is a local object, determined by information near pp alone. To state this precisely we first set up a framework for “functions defined only near pp”.

Definition 2.1Germ of a function

Let pMp \in M. On the collection of all pairs (U,f)(U, f), where UU is an open set containing pp and fC(U)f \in C^{\infty}(U), define a relation \sim by

(U,f)(V,g)    there is an open WUV with pW and fW=gW.(U, f) \sim (V, g) \iff \text{there is an open $W \subset U \cap V$ with $p \in W$ and } f|_W = g|_W .

Then \sim is an equivalence relation; an equivalence class is called a germ at pp of a CC^{\infty} function, written [f]p[f]_p or simply ff. The set of all germs is denoted Cp(M)C^{\infty}_p(M).

Let us check that \sim is an equivalence relation. Reflexivity and symmetry are of an obvious shape, but written out: taking W=UW = U gives (U,f)(U,f)(U,f) \sim (U,f), and symmetry follows because the definition is symmetric in ff and gg. For transitivity, if W1W_1 witnesses (U,f)(V,g)(U,f) \sim (V,g) and W2W_2 witnesses (V,g)(V,h)(V,g) \sim (V', h), then W1W2W_1 \cap W_2 is an open set containing pp on which f=g=hf = g = h.

Sums, products and scalar multiples are defined on Cp(M)C^{\infty}_p(M) representative by representative, on the intersection of the domains (independence of the representatives follows by intersecting the sets WW above), and this makes Cp(M)C^{\infty}_p(M) a commutative algebra over R\mathbb{R}. Moreover [f]pf(p)[f]_p \mapsto f(p) is a well-defined algebra homomorphism Cp(M)RC^{\infty}_p(M) \to \mathbb{R}.

3. The first definition: the velocity of a curve

Section titled “3. The first definition: the velocity of a curve”

We begin with the most geometric definition. A curve through pp is a CC^{\infty} map γ:(ε,ε)M\gamma : (-\varepsilon, \varepsilon) \to M with ε>0\varepsilon > 0 and γ(0)=p\gamma(0) = p. We define what it means for two curves to “have the same velocity at pp” using a chart. Of course this is a definition only once we have checked that it does not depend on the chart.

Lemma 3.1Agreement of velocities is chart-independent

Let pMp \in M and let γ1,γ2\gamma_1, \gamma_2 be CC^{\infty} curves through pp. If (U,φ)(U, \varphi) and (V,ψ)(V, \psi) are charts both containing pp, then

(φγ1)(0)=(φγ2)(0)    (ψγ1)(0)=(ψγ2)(0).(\varphi \circ \gamma_1)'(0) = (\varphi \circ \gamma_2)'(0) \iff (\psi \circ \gamma_1)'(0) = (\psi \circ \gamma_2)'(0).
Proof(Lemma 3.1)

Since each γj\gamma_j is continuous, after shrinking ε\varepsilon we may assume γj((ε,ε))UV\gamma_j((-\varepsilon,\varepsilon)) \subset U \cap V. The map τ=ψφ1\tau = \psi \circ \varphi^{-1} is a CC^{\infty} diffeomorphism on φ(UV)\varphi(U \cap V) (this is the compatibility condition in the atlas), and

ψγj=(ψφ1)(φγj)=τ(φγj).\psi \circ \gamma_j = (\psi \circ \varphi^{-1}) \circ (\varphi \circ \gamma_j) = \tau \circ (\varphi \circ \gamma_j).

Since φγj:(ε,ε)Rn\varphi \circ \gamma_j : (-\varepsilon,\varepsilon) \to \mathbb{R}^n is CC^{\infty}, the several-variable chain rule(Theorem 6.1)[多変数関数の微分と偏微分] gives

(ψγj)(0)=Dτ(φ(p))(φγj)(0).(\psi \circ \gamma_j)'(0) = D\tau\big(\varphi(p)\big)\,(\varphi \circ \gamma_j)'(0).

Here Dτ(φ(p))D\tau(\varphi(p)) is the Jacobian matrix of τ\tau, and it is invertible because τ\tau is a diffeomorphism. Multiplying by an invertible matrix is injective, so (φγ1)(0)=(φγ2)(0)(\varphi\circ\gamma_1)'(0) = (\varphi\circ\gamma_2)'(0) and (ψγ1)(0)=(ψγ2)(0)(\psi\circ\gamma_1)'(0) = (\psi\circ\gamma_2)'(0) are equivalent.

Definition 3.2Tangent vector (definition by curves)

Write Cp\mathcal{C}_p for the set of all CC^{\infty} curves through pp. Declare γ1cγ2\gamma_1 \sim_{\mathrm{c}} \gamma_2 when (φγ1)(0)=(φγ2)(0)(\varphi\circ\gamma_1)'(0) = (\varphi\circ\gamma_2)'(0) for some chart (U,φ)(U,\varphi) containing pp. By Lemma 3.1 this does not depend on the chart chosen, and it is evidently an equivalence relation. The quotient Cp/c\mathcal{C}_p/\sim_{\mathrm{c}} is denoted TpcMT^{\mathrm{c}}_pM, and its elements [γ][\gamma] are called tangent vectors at pp.

The virtue of this definition is that intuition works. Its defect is that sums and scalar multiples are not available on the spot: there is no operation on a manifold that “adds” two curves. To define the sum one has to choose a chart, add the components there, and pass back to a curve — and then check separately that the result is independent of the chart. This inconvenience is precisely the motivation for the next two definitions.

4. The second definition: coordinate components and the transformation law

Section titled “4. The second definition: coordinate components and the transformation law”

We now rewrite the classical style of tensor analysis in modern language. The definition reads: a vector is a rule assigning nn numbers to each coordinate system, in such a way that a change of coordinates transforms them by the Jacobian matrix.

Let Ap\mathcal{A}_p be the set of all pairs (φ,v)(\varphi, v) consisting of a chart (U,φ)(U,\varphi) containing pp and a vector vRnv \in \mathbb{R}^n.

Definition 4.1Tangent vector (definition by coordinate components)

For (φ,v),(ψ,w)Ap(\varphi, v), (\psi, w) \in \mathcal{A}_p set

(φ,v)o(ψ,w)    w=D(ψφ1)(φ(p))v.(\varphi, v) \sim_{\mathrm{o}} (\psi, w) \iff w = D(\psi \circ \varphi^{-1})\big(\varphi(p)\big)\, v .

The quotient set Ap/o\mathcal{A}_p/\sim_{\mathrm{o}} is denoted TpoMT^{\mathrm{o}}_pM.

Proposition 4.2The transformation law defines an equivalence relation

o\sim_{\mathrm{o}} is an equivalence relation on Ap\mathcal{A}_p.

Proof(Proposition 4.2)

To shorten the notation write Jψφ=D(ψφ1)(φ(p))J_{\psi\varphi} = D(\psi\circ\varphi^{-1})(\varphi(p)).

Reflexivity. φφ1\varphi \circ \varphi^{-1} is the identity on φ(U)\varphi(U), so Jφφ=IJ_{\varphi\varphi} = I (the identity matrix) and v=Ivv = Iv.

Symmetry. Since (ψφ1)1=φψ1(\psi\circ\varphi^{-1})^{-1} = \varphi\circ\psi^{-1}, the formula for the derivative of an inverse map on Rn\mathbb{R}^n (apply the chain rule to φψ1ψφ1=id\varphi\psi^{-1}\circ\psi\varphi^{-1} = \mathrm{id}) gives Jφψ=Jψφ1J_{\varphi\psi} = J_{\psi\varphi}^{-1}. Hence w=Jψφvw = J_{\psi\varphi}v implies v=Jφψwv = J_{\varphi\psi} w.

Transitivity. For a third chart (W,χ)(W,\chi) we have χφ1=(χψ1)(ψφ1)\chi\circ\varphi^{-1} = (\chi\circ\psi^{-1})\circ(\psi\circ\varphi^{-1}) on φ(UVW)\varphi(U\cap V\cap W), so the chain rule yields Jχφ=JχψJψφJ_{\chi\varphi} = J_{\chi\psi}J_{\psi\varphi}. Therefore w=Jψφvw = J_{\psi\varphi}v and u=Jχψwu = J_{\chi\psi}w imply u=Jχφvu = J_{\chi\varphi}v.

Note that all three properties follow from identities valid on intersections of chart domains, and those intersections are open sets containing pp, hence nonempty.

The set TpoMT^{\mathrm{o}}_pM carries a natural vector space structure. Fixing one chart φ\varphi, the map v[(φ,v)]v \mapsto [(\varphi, v)] is a bijection RnTpoM\mathbb{R}^n \to T^{\mathrm{o}}_pM (surjective by definition, injective because Jφφ=IJ_{\varphi\varphi} = I), and we transport the linear structure of Rn\mathbb{R}^n through it. The transported structure is independent of the chart because the transition vJψφvv \mapsto J_{\psi\varphi}v defining o\sim_{\mathrm{o}} is linear. This “gluing by linear maps” is also the reason why the tangent bundle will turn out to be a vector bundle.

Remark 4.3Correspondence with the index notation of physics

The sentence found in textbooks of general relativity — “a contravariant vector is a quantity that transforms under a change of coordinates as Vμ=xμxνVνV'^{\mu} = \dfrac{\partial x'^{\mu}}{\partial x^{\nu}} V^{\nu}” — is nothing but o\sim_{\mathrm{o}} of Definition 4.1 written with the Einstein summation convention. Here xμxν\dfrac{\partial x'^{\mu}}{\partial x^{\nu}} is the (μ,ν)(\mu,\nu) entry of JψφJ_{\psi\varphi}. A covariant vector (a 11-form) is a quantity transforming by the inverse matrix, which is to say an element of the dual space of TpMT_pM.

The third definition regards a tangent vector as “an operator taking directional derivatives of functions”. Recall the directional derivative Dvf(a)=iviif(a)D_v f(a) = \sum_i v^i \partial_i f(a) on Rn\mathbb{R}^n: it is a real-valued linear functional fDvf(a)f \mapsto D_vf(a) satisfying the product rule. Conversely, we shall prove that any functional satisfying those two conditions is a directional derivative.

Definition 5.1Derivation at a point

An R\mathbb{R}-linear map X:Cp(M)RX : C^{\infty}_p(M) \to \mathbb{R} satisfying

X(fg)=X(f)g(p)+f(p)X(g)X(fg) = X(f)\,g(p) + f(p)\,X(g)

for all f,gCp(M)f, g \in C^{\infty}_p(M) (the Leibniz rule) is called a derivation at pp. The set of all derivations is denoted TpdMT^{\mathrm{d}}_pM.

That TpdMT^{\mathrm{d}}_pM becomes a real vector space under (X+Y)(f)=X(f)+Y(f)(X+Y)(f) = X(f)+Y(f) and (cX)(f)=cX(f)(cX)(f) = cX(f) is immediate, since both defining conditions behave linearly. Indeed, (X+Y)(fg)=X(fg)+Y(fg)=(X(f)+Y(f))g(p)+f(p)(X(g)+Y(g))(X+Y)(fg) = X(fg)+Y(fg) = (X(f)+Y(f))g(p) + f(p)(X(g)+Y(g)).

Lemma 5.2Derivations kill constants

Let XTpdMX \in T^{\mathrm{d}}_pM and, for cRc \in \mathbb{R}, write c\underline{c} for the germ of the constant function cc. Then X(c)=0X(\underline{c}) = 0.

Proof(Lemma 5.2)

Applying the Leibniz rule to 11=1\underline{1} \cdot \underline{1} = \underline{1} gives

X(1)=X(1)1+1X(1)=2X(1),X(\underline{1}) = X(\underline{1})\cdot 1 + 1 \cdot X(\underline{1}) = 2X(\underline{1}),

so X(1)=0X(\underline{1}) = 0. For general cc, write c=c1\underline{c} = c\,\underline{1} and use linearity: X(c)=cX(1)=0X(\underline{c}) = cX(\underline{1}) = 0.

The next result is the analytic lemma at the core of the theory. Think of it as the version of Taylor’s theorem in which the remainder is written as a smooth function.

Lemma 5.3Hadamard's lemma

Let ARnA \subset \mathbb{R}^n be an open set that is star-shaped about aAa \in A (that is, for every xAx \in A the segment {a+t(xa):t[0,1]}\{a + t(x-a) : t \in [0,1]\} lies in AA), and let hC(A)h \in C^{\infty}(A). Then there exist g1,,gnC(A)g_1, \ldots, g_n \in C^{\infty}(A) such that for all xAx \in A

h(x)=h(a)+i=1n(xiai)gi(x),gi(a)=ih(a).h(x) = h(a) + \sum_{i=1}^{n} (x^i - a^i)\, g_i(x), \qquad g_i(a) = \partial_i h(a).
Proof(Lemma 5.3)

Fix xAx \in A and put u(t)=h(a+t(xa))u(t) = h\big(a + t(x-a)\big). Star-shapedness gives a+t(xa)Aa+t(x-a) \in A for t[0,1]t \in [0,1], so uu is defined on [0,1][0,1]; it is CC^{\infty} because hh is and by the chain rule, and

u(t)=i=1n(xiai)ih(a+t(xa)).u'(t) = \sum_{i=1}^{n} (x^i - a^i)\,\partial_i h\big(a+t(x-a)\big).

By the fundamental theorem of calculus(Theorem 5.4)[積分の基本定理と定積分] we have u(1)u(0)=01u(t)dtu(1) - u(0) = \int_0^1 u'(t)\,dt, that is,

h(x)h(a)=i=1n(xiai)01ih(a+t(xa))dt.h(x) - h(a) = \sum_{i=1}^{n} (x^i - a^i) \int_0^1 \partial_i h\big(a+t(x-a)\big)\,dt.

So setting

gi(x)=01ih(a+t(xa))dtg_i(x) = \int_0^1 \partial_i h\big(a + t(x-a)\big)\,dt

gives the desired decomposition. Each gig_i is CC^{\infty} because the integrand is a CC^{\infty} function of (t,x)[0,1]×A(t,x) \in [0,1]\times A and the interval [0,1][0,1] is compact, so differentiation under the integral sign is legitimate to every order. Finally, substituting x=ax = a makes the integrand the constant ih(a)\partial_i h(a), whence gi(a)=ih(a)g_i(a) = \partial_i h(a).

Star-shapedness is essential. When we use this lemma on a manifold we first replace the image of the chart by an open ball centered at φ(p)\varphi(p) (shrinking the domain loses no information, since we are dealing with germs).

Theorem 5.4A basis for the space of derivations

Let pMp \in M, let (U,φ)(U,\varphi) be a chart containing pp with φ=(x1,,xn)\varphi = (x^1,\ldots,x^n), and put a=φ(p)a = \varphi(p). For each ii define

xip(f):=i(fφ1)(a)(fCp(M)).\left.\frac{\partial}{\partial x^i}\right|_p (f) := \partial_i\big(f \circ \varphi^{-1}\big)(a) \qquad (f \in C^{\infty}_p(M)).

Then /xipTpdM\partial/\partial x^i|_p \in T^{\mathrm{d}}_pM, and (/x1p,,/xnp)\left(\partial/\partial x^1|_p, \ldots, \partial/\partial x^n|_p\right) is a basis of TpdMT^{\mathrm{d}}_pM. In particular dimTpdM=n\dim T^{\mathrm{d}}_pM = n, and every XTpdMX \in T^{\mathrm{d}}_pM is uniquely expressed as

X=i=1nX(xi)xip.X = \sum_{i=1}^{n} X(x^i)\left.\frac{\partial}{\partial x^i}\right|_p .
Proof(Theorem 5.4)

First we check that the right-hand side does not depend on the representative of the germ. If ff and f~\tilde f agree on a neighborhood WW of pp, then fφ1f\circ\varphi^{-1} and f~φ1\tilde f\circ\varphi^{-1} agree on φ(WU)\varphi(W\cap U), an open neighborhood of aa, so their partial derivatives at aa coincide.

(1) It is a derivation. Linearity follows from linearity of partial differentiation. For the Leibniz rule, use (fg)φ1=(fφ1)(gφ1)(fg)\circ\varphi^{-1} = (f\circ\varphi^{-1})\cdot(g\circ\varphi^{-1}) and the product rule on Rn\mathbb{R}^n:

i((fg)φ1)(a)=i(fφ1)(a)(gφ1)(a)+(fφ1)(a)i(gφ1)(a),\partial_i\big((fg)\circ\varphi^{-1}\big)(a) = \partial_i(f\circ\varphi^{-1})(a)\,(g\circ\varphi^{-1})(a) + (f\circ\varphi^{-1})(a)\,\partial_i(g\circ\varphi^{-1})(a),

and since (fφ1)(a)=f(p)(f\circ\varphi^{-1})(a) = f(p) and (gφ1)(a)=g(p)(g\circ\varphi^{-1})(a) = g(p), this is the asserted identity.

(2) Linear independence. Suppose ici/xip=0\sum_i c^i \partial/\partial x^i|_p = 0 and apply it to the germ of the coordinate function xjC(U)x^j \in C^{\infty}(U). Since xjφ1x^j \circ \varphi^{-1} is the jj-th coordinate function on φ(U)\varphi(U), we have i(xjφ1)=δij\partial_i(x^j\circ\varphi^{-1}) = \delta^j_i (the Kronecker delta). Hence

0=icii(xjφ1)(a)=iciδij=cj0 = \sum_i c^i\,\partial_i(x^j\circ\varphi^{-1})(a) = \sum_i c^i \delta^j_i = c^j

for every jj, so c1==cn=0c^1 = \cdots = c^n = 0.

(3) They span. Take XTpdMX \in T^{\mathrm{d}}_pM and set ci=X(xi)c^i = X(x^i). Let fCp(M)f \in C^{\infty}_p(M) be arbitrary; shrinking the domain of a representative, we may assume that φ(U)\varphi(U) is an open ball BB centered at aa (a ball is star-shaped). Applying Lemma 5.3 to h=fφ1C(B)h = f\circ\varphi^{-1} \in C^{\infty}(B) produces giC(B)g_i \in C^{\infty}(B) with

h=h(a)+i(i-th coordinateai)gion B,gi(a)=ih(a).h = h(a) + \sum_i (\text{$i$-th coordinate} - a^i)\,g_i \quad\text{on } B, \qquad g_i(a) = \partial_i h(a).

Composing both sides with φ\varphi gives the identity

f=f(p)+i=1n(xiai)(giφ)f = \underline{f(p)} + \sum_{i=1}^{n} (x^i - \underline{a^i})\,(g_i\circ\varphi)

on UU. Read this as an identity of germs and apply XX. Linearity, Lemma 5.2 and the Leibniz rule give

X(f)=X(f(p))+i[X(xiai)(giφ)(p)+(xi(p)ai)X(giφ)]=0+i[(X(xi)0)gi(a)+0]=iciih(a)=icixip(f).\begin{aligned} X(f) &= X(\underline{f(p)}) + \sum_i \Big[ X(x^i - \underline{a^i})\cdot (g_i\circ\varphi)(p) + \big(x^i(p) - a^i\big)\cdot X(g_i\circ\varphi) \Big] \\ &= 0 + \sum_i \Big[ \big(X(x^i) - 0\big)\, g_i(a) + 0 \Big] = \sum_i c^i\, \partial_i h(a) = \sum_i c^i \left.\frac{\partial}{\partial x^i}\right|_p (f). \end{aligned}

In the second line the second term was killed using xi(p)=aix^i(p) = a^i (the ii-th component of φ(p)=a\varphi(p) = a). As ff was arbitrary, X=ici/xipX = \sum_i c^i \partial/\partial x^i|_p.

From (2) and (3) we get a basis, and uniqueness of the coefficients together with the computation in (2) pins them down as ci=X(xi)c^i = X(x^i).

Remark 5.5Why CC^{\infty} is indispensable

The proof of Theorem 5.4 rests entirely on Lemma 5.3, and that lemma uses smoothness of hh to obtain smoothness of the gig_i. If one takes the same definition of derivation on the algebra of germs of CkC^k functions with kk finite, the gig_i are only guaranteed to be Ck1C^{k-1} and the proof collapses. The collapse is genuine: the space of derivations on germs of CkC^k functions is infinite-dimensional. See the Appendix for details.

Before going on, let us confirm that defining derivations on the global algebra of functions, without germs, gives the same object. This identification is used constantly in practice.

Proposition 5.6Global derivations agree with derivations on germs

Let T~pM\widetilde{T}_pM be the set of all R\mathbb{R}-linear maps X:C(M)RX : C^{\infty}(M) \to \mathbb{R} satisfying the Leibniz rule X(fg)=X(f)g(p)+f(p)X(g)X(fg) = X(f)g(p) + f(p)X(g). Then the restriction map

Φ:TpdMT~pM,Φ(Y)(f)=Y([f]p)\Phi : T^{\mathrm{d}}_pM \to \widetilde{T}_pM, \qquad \Phi(Y)(f) = Y([f]_p)

is a linear isomorphism.

Proof(Proposition 5.6)

On MM there exist bump functions: for any open neighborhood VV of pp one can find χC(M)\chi \in C^{\infty}(M) with χ1\chi \equiv 1 on some neighborhood of pp and suppχV\operatorname{supp}\chi \subset V (this is the partition-of-unity argument in Smooth manifolds).

Linearity of Φ\Phi is clear from the definition: Φ(Y1+cY2)(f)=(Y1+cY2)([f]p)=Φ(Y1)(f)+cΦ(Y2)(f)\Phi(Y_1+cY_2)(f) = (Y_1+cY_2)([f]_p) = \Phi(Y_1)(f) + c\Phi(Y_2)(f). And Φ(Y)\Phi(Y) satisfies the Leibniz rule because [fg]p=[f]p[g]p[fg]_p = [f]_p[g]_p.

Injectivity. Suppose Φ(Y)=0\Phi(Y) = 0. Given a germ [f]p[f]_p, choose a representative fC(V)f \in C^{\infty}(V) and, with χ\chi as above, set f^=χf\hat f = \chi f (defined to be 00 outside VV). Then f^C(M)\hat f \in C^{\infty}(M) and f^=f\hat f = f on a neighborhood of pp, so [f^]p=[f]p[\hat f]_p = [f]_p. Hence Y([f]p)=Y([f^]p)=Φ(Y)(f^)=0Y([f]_p) = Y([\hat f]_p) = \Phi(Y)(\hat f) = 0, and therefore Y=0Y = 0.

Surjectivity. Let XT~pMX \in \widetilde{T}_pM. We first prove locality. Suppose u,vC(M)u, v \in C^{\infty}(M) agree on a neighborhood WW of pp and put w=uvw = u - v, so that wW=0w|_W = 0. Choose a bump function χ\chi with suppχW\operatorname{supp}\chi \subset W and χ1\chi\equiv 1 near pp; then χw0\chi w \equiv 0 (on WW we have w=0w = 0, outside WW we have χ=0\chi = 0). Since X(0)=0X(0) = 0 by linearity, the Leibniz rule gives

0=X(χw)=X(χ)w(p)+χ(p)X(w)=0+X(w)0 = X(\chi w) = X(\chi)\,w(p) + \chi(p)\,X(w) = 0 + X(w)

(using w(p)=0w(p) = 0 and χ(p)=1\chi(p) = 1). Hence X(u)=X(v)X(u) = X(v).

Now define Y([f]p):=X(χf)Y([f]_p) := X(\chi f), where χ\chi is a bump function supported in the domain of the representative. By locality this depends on neither the representative nor χ\chi. Linearity is clear. For the Leibniz rule, note that χ(fg)\chi(fg) and (χf)(χg)(\chi f)(\chi g) agree on the neighborhood where χ1\chi \equiv 1, so

Y([f]p[g]p)=X(χfg)=X((χf)(χg))=X(χf)g(p)+f(p)X(χg)=Y([f]p)g(p)+f(p)Y([g]p).Y([f]_p[g]_p) = X\big(\chi fg\big) = X\big((\chi f)(\chi g)\big) = X(\chi f)g(p) + f(p)X(\chi g) = Y([f]_p)g(p) + f(p)Y([g]_p).

Finally Φ(Y)(f)=Y([f]p)=X(χf)=X(f)\Phi(Y)(f) = Y([f]_p) = X(\chi f) = X(f) by locality, so Φ(Y)=X\Phi(Y) = X.

All the players are now on stage. We show that the three constructions are naturally isomorphic, “naturally” meaning that the maps giving the isomorphisms do not depend on a choice of chart.

flowchart LR
A["Equivalence classes of curves T_p^c M<br/>curves carrying a velocity"] -->|"Λ: send a curve to its components (φ∘γ)'(0)"| B["Coordinate components T_p^o M<br/>n numbers transforming by the Jacobian"]
B -->|"Θ: send components v to Σ vⁱ ∂/∂xⁱ"| C["Derivations T_p^d M<br/>functionals satisfying the Leibniz rule"]
A -->|"Ξ: send a curve to f ↦ (f∘γ)'(0)"| C
The three definitions and the natural maps linking them; the triangle commutes.

Theorem 6.1Equivalence of the three tangent spaces

Let pMp \in M. Each of the following maps is a bijection, and Ξ=ΘΛ\Xi = \Theta \circ \Lambda.

  1. Λ:TpcMTpoM\Lambda : T^{\mathrm{c}}_pM \to T^{\mathrm{o}}_pM, Λ([γ])=[(φ,(φγ)(0))]\Lambda([\gamma]) = \big[(\varphi, (\varphi\circ\gamma)'(0))\big], where (U,φ)(U,\varphi) is any chart containing pp.
  2. Θ:TpoMTpdM\Theta : T^{\mathrm{o}}_pM \to T^{\mathrm{d}}_pM, Θ([(φ,v)])=i=1nvixip\Theta\big([(\varphi,v)]\big) = \sum_{i=1}^n v^i \left.\dfrac{\partial}{\partial x^i}\right|_p.
  3. Ξ:TpcMTpdM\Xi : T^{\mathrm{c}}_pM \to T^{\mathrm{d}}_pM, Ξ([γ])(f)=(fγ)(0)\Xi([\gamma])(f) = (f\circ\gamma)'(0).

Moreover Θ\Theta is a linear isomorphism, and if TpcMT^{\mathrm{c}}_pM is given the linear structure transported by Ξ\Xi, then all three are isomorphic as vector spaces. From now on we identify them and write TpMT_pM.

Proof(Theorem 6.1)

Λ\Lambda is well defined and bijective. If γ1cγ2\gamma_1 \sim_{\mathrm{c}} \gamma_2, then by definition their velocities agree in the same chart, so the value of Λ\Lambda is independent of the representative. Independence of the chart follows from the identity (ψγ)(0)=Jψφ(φγ)(0)(\psi\circ\gamma)'(0) = J_{\psi\varphi}(\varphi\circ\gamma)'(0) established in the proof of Lemma 3.1: it says exactly that (φ,(φγ)(0))o(ψ,(ψγ)(0))(\varphi, (\varphi\circ\gamma)'(0)) \sim_{\mathrm{o}} (\psi, (\psi\circ\gamma)'(0)). Injectivity: reading Λ([γ1])=Λ([γ2])\Lambda([\gamma_1]) = \Lambda([\gamma_2]) through the representative in the chart φ\varphi gives (φγ1)(0)=(φγ2)(0)(\varphi\circ\gamma_1)'(0) = (\varphi\circ\gamma_2)'(0), i.e. γ1cγ2\gamma_1\sim_{\mathrm{c}}\gamma_2. Surjectivity: given [(φ,v)][(\varphi,v)], put a=φ(p)a = \varphi(p) and define

γ(t)=φ1(a+tv)(t<ε).\gamma(t) = \varphi^{-1}(a + tv) \qquad (|t| < \varepsilon).

Since φ(U)\varphi(U) is open, ε>0\varepsilon > 0 can be taken small enough that a+tvφ(U)a+tv \in \varphi(U); then γ\gamma is CC^{\infty} as a composition of CC^{\infty} maps, γ(0)=p\gamma(0) = p, and (φγ)(t)=a+tv(\varphi\circ\gamma)(t) = a+tv gives (φγ)(0)=v(\varphi\circ\gamma)'(0) = v.

Θ\Theta is well defined. Suppose (φ,v)o(ψ,w)(\varphi,v)\sim_{\mathrm{o}}(\psi,w), i.e. w=Jψφvw = J_{\psi\varphi}v. Write ψ=(y1,,yn)\psi = (y^1,\ldots,y^n), τ=ψφ1\tau = \psi\circ\varphi^{-1} and a=φ(p)a = \varphi(p), so that the (j,i)(j,i) entry of JψφJ_{\psi\varphi} is iτj(a)\partial_i\tau^j(a). For any germ ff, applying the chain rule on Rn\mathbb{R}^n to fφ1=(fψ1)τf\circ\varphi^{-1} = (f\circ\psi^{-1})\circ\tau gives

xip(f)=i((fψ1)τ)(a)=j=1nj(fψ1)(τ(a))iτj(a)=j=1niτj(a)yjp(f)\left.\frac{\partial}{\partial x^i}\right|_p (f) = \partial_i\big((f\circ\psi^{-1})\circ\tau\big)(a) = \sum_{j=1}^n \partial_j(f\circ\psi^{-1})\big(\tau(a)\big)\,\partial_i\tau^j(a) = \sum_{j=1}^n \partial_i\tau^j(a) \left.\frac{\partial}{\partial y^j}\right|_p(f)

(using τ(a)=ψ(p)\tau(a) = \psi(p)). Therefore

ivixip=j(iiτj(a)vi)yjp=jwjyjp,\sum_i v^i \left.\frac{\partial}{\partial x^i}\right|_p = \sum_j \Big(\sum_i \partial_i\tau^j(a)\,v^i\Big) \left.\frac{\partial}{\partial y^j}\right|_p = \sum_j w^j \left.\frac{\partial}{\partial y^j}\right|_p ,

so the value of Θ\Theta is independent of the representative.

Θ\Theta is a linear isomorphism. Fixing a chart φ\varphi, the correspondence TpoM[(φ,v)]vRnT^{\mathrm{o}}_pM \ni [(\varphi,v)] \leftrightarrow v \in \mathbb{R}^n is a linear isomorphism (this is exactly how the linear structure was defined in §4), and under it Θ\Theta becomes vivi/xipv \mapsto \sum_i v^i \partial/\partial x^i|_p. By Theorem 5.4 the family {/xip}\{\partial/\partial x^i|_p\} is a basis of TpdMT^{\mathrm{d}}_pM, so this map is a linear isomorphism.

Ξ=ΘΛ\Xi = \Theta\circ\Lambda. Let γ\gamma be a curve through pp and v=(φγ)(0)v = (\varphi\circ\gamma)'(0). For any germ ff we have fγ=(fφ1)(φγ)f\circ\gamma = (f\circ\varphi^{-1})\circ(\varphi\circ\gamma), so the chain rule gives

(fγ)(0)=ii(fφ1)(φ(p))((φγ)i)(0)=ivixip(f).(f\circ\gamma)'(0) = \sum_i \partial_i(f\circ\varphi^{-1})\big(\varphi(p)\big)\cdot \big((\varphi\circ\gamma)^i\big)'(0) = \sum_i v^i \left.\frac{\partial}{\partial x^i}\right|_p(f).

The right-hand side is Θ(Λ([γ]))(f)\Theta(\Lambda([\gamma]))(f). Incidentally this also shows that Ξ([γ])\Xi([\gamma]) is a derivation, being in the image of ΘΛ\Theta\circ\Lambda. That Ξ\Xi is a bijection follows, as a composition, from the bijectivity of Λ\Lambda and Θ\Theta.

Example 6.2The case M=RnM = \mathbb{R}^n

Equip M=RnM = \mathbb{R}^n with the atlas whose only chart is the identity map id\mathrm{id}. The coordinate functions are the usual xix^i, and /xia\partial/\partial x^i|_a is the ordinary partial derivative fif(a)f \mapsto \partial_i f(a). By Theorem 5.4, TaRnT_a\mathbb{R}^n is nn-dimensional, and the map sending vRnv \in \mathbb{R}^n to ivi/xia\sum_i v^i\partial/\partial x^i|_a is a linear isomorphism Rn  TaRn\mathbb{R}^n \xrightarrow{\ \sim\ } T_a\mathbb{R}^n. This isomorphism is canonical — no chart has to be chosen — so from now on we identify TaRnT_a\mathbb{R}^n with Rn\mathbb{R}^n. On the side of curves it is the familiar correspondence [γ]γ(0)[\gamma] \mapsto \gamma'(0).

Example 6.3Rewriting the basis in polar coordinates

Take M=R2M = \mathbb{R}^2 and U={(x,y):x>0}U = \{(x,y) : x > 0\}, and put the polar chart ψ=(r,θ)\psi = (r,\theta) on UU by prescribing its inverse

ψ1(r,θ)=(rcosθ, rsinθ),r>0, π2<θ<π2.\psi^{-1}(r,\theta) = (r\cos\theta,\ r\sin\theta), \qquad r > 0,\ -\tfrac{\pi}{2} < \theta < \tfrac{\pi}{2}.

We use the transformation formula from the proof of Theorem 6.1, with polar coordinates as the source and Cartesian coordinates as the target. The Jacobian matrix of τ=(Cartesian)(polar)1\tau = (\text{Cartesian})\circ(\text{polar})^{-1} is

Dτ(r,θ)=(cosθrsinθsinθrcosθ),D\tau(r,\theta) = \begin{pmatrix} \cos\theta & -r\sin\theta \\ \sin\theta & r\cos\theta \end{pmatrix},

so that

rp=cosθxp+sinθyp,θp=rsinθxp+rcosθyp.\left.\frac{\partial}{\partial r}\right|_p = \cos\theta \left.\frac{\partial}{\partial x}\right|_p + \sin\theta \left.\frac{\partial}{\partial y}\right|_p, \qquad \left.\frac{\partial}{\partial \theta}\right|_p = -r\sin\theta \left.\frac{\partial}{\partial x}\right|_p + r\cos\theta \left.\frac{\partial}{\partial y}\right|_p.

Let us compute concretely at p=(1,1)p = (1,1), that is, r=2r = \sqrt{2} and θ=π/4\theta = \pi/4. Since cosθ=sinθ=1/2\cos\theta = \sin\theta = 1/\sqrt{2},

rp=12(xp+yp),θp=xp+yp.\left.\frac{\partial}{\partial r}\right|_p = \frac{1}{\sqrt2}\left(\left.\frac{\partial}{\partial x}\right|_p + \left.\frac{\partial}{\partial y}\right|_p\right), \qquad \left.\frac{\partial}{\partial\theta}\right|_p = -\left.\frac{\partial}{\partial x}\right|_p + \left.\frac{\partial}{\partial y}\right|_p .

As a check, take f(x,y)=x2+y2f(x,y) = x^2+y^2. In polar coordinates fψ1(r,θ)=r2f\circ\psi^{-1}(r,\theta) = r^2, so /rp(f)=2r=22\partial/\partial r|_p (f) = 2r = 2\sqrt2 and /θp(f)=0\partial/\partial\theta|_p(f) = 0. Computing instead with the right-hand sides,

12(2x+2y)(1,1)=42=22,(2x+2y)(1,1)=0,\frac{1}{\sqrt2}\big(2x + 2y\big)\Big|_{(1,1)} = \frac{4}{\sqrt2} = 2\sqrt2, \qquad \big(-2x + 2y\big)\Big|_{(1,1)} = 0,

in agreement. Note that /θp\partial/\partial\theta|_p is not a “unit vector”. We have not yet introduced any notion of length, but its Cartesian components are (1,1)(-1,1), of Euclidean length 2=r\sqrt2 = r. That a coordinate basis need not be orthonormal is the first thing to get used to when working with curvilinear coordinates.

The greatest dividend of introducing tangent spaces is that a smooth map can be linearized at each point. With the definition by derivations, the differential can be written down with astonishing brevity.

Definition 7.1The differential of a smooth map

Let F:MNF : M \to N be a CC^{\infty} map, pMp \in M and q=F(p)q = F(p). For XTpMX \in T_pM define dFp(X)dF_p(X) by

(dFp(X))(g):=X(gF)(gCq(N)).\big(dF_p(X)\big)(g) := X(g\circ F) \qquad \big(g \in C^{\infty}_q(N)\big).

The map dFp:TpMTqNdF_p : T_pM \to T_qN is called the differential of FF at pp (also the pushforward, written FpF_{*p}).

Let us first check that the definition makes sense. If gg is a CC^{\infty} function on a neighborhood VV of qq, then F1(V)F^{-1}(V) is an open neighborhood of pp by continuity of FF, and gFC(F1(V))g\circ F \in C^{\infty}(F^{-1}(V)), so [gF]pCp(M)[g\circ F]_p \in C^{\infty}_p(M) is well defined. Replacing the representative of gg on a neighborhood of qq does not change gFg\circ F on a neighborhood of pp, so we obtain a map of germs F:Cq(N)Cp(M)F^{*} : C^{\infty}_q(N) \to C^{\infty}_p(M), [g][gF][g] \mapsto [g\circ F], and it is an algebra homomorphism.

Proposition 7.2The differential is a linear map

In Definition 7.1, dFp(X)dF_p(X) is a derivation at qq, and dFp:TpMTqNdF_p : T_pM \to T_qN is linear.

Proof(Proposition 7.2)

Linearity of dFp(X)dF_p(X) holds because FF^{*} is linear and XX is linear. For the Leibniz rule, use that FF^{*} preserves products, (gh)F=(gF)(hF)(gh)\circ F = (g\circ F)(h\circ F), together with the Leibniz rule for XX:

dFp(X)(gh)=X((gF)(hF))=X(gF)h(F(p))+g(F(p))X(hF)=dFp(X)(g)h(q)+g(q)dFp(X)(h)dF_p(X)(gh) = X\big((g\circ F)(h\circ F)\big) = X(g\circ F)\,h(F(p)) + g(F(p))\,X(h\circ F) = dF_p(X)(g)\,h(q) + g(q)\,dF_p(X)(h)

(where we used (gF)(p)=g(q)(g\circ F)(p) = g(q)). Linearity of XdFp(X)X \mapsto dF_p(X) is immediate from dFp(X+cY)(g)=(X+cY)(gF)=X(gF)+cY(gF)dF_p(X+cY)(g) = (X+cY)(g\circ F) = X(g\circ F) + cY(g\circ F).

Proposition 7.3Description in terms of curves

In the situation above, if X=Ξ([γ])TpMX = \Xi([\gamma]) \in T_pM then dFp(X)=Ξ([Fγ])dF_p(X) = \Xi([F\circ\gamma]). In the language of curves, then, the differential is nothing but the operation of pushing a curve forward by FF.

Proof(Proposition 7.3)

Being a composition of CC^{\infty} maps, FγF\circ\gamma is a CC^{\infty} curve through qq. For any gCq(N)g \in C^{\infty}_q(N),

Ξ([Fγ])(g)=(g(Fγ))(0)=((gF)γ)(0)=Ξ([γ])(gF)=X(gF)=dFp(X)(g).\Xi([F\circ\gamma])(g) = \big(g\circ (F\circ\gamma)\big)'(0) = \big((g\circ F)\circ\gamma\big)'(0) = \Xi([\gamma])(g\circ F) = X(g\circ F) = dF_p(X)(g).

Nothing beyond associativity of composition has been used.

Theorem 7.4The chain rule

Let F:MNF : M \to N and G:NPG : N \to P be CC^{\infty} maps and pMp \in M. Then

d(GF)p=dGF(p)dFp:TpMTG(F(p))P.d(G\circ F)_p = dG_{F(p)} \circ dF_p : T_pM \to T_{G(F(p))}P .

Moreover d(idM)p=idTpMd(\mathrm{id}_M)_p = \mathrm{id}_{T_pM}.

Proof(Theorem 7.4)

Let XTpMX \in T_pM and hCG(F(p))(P)h \in C^{\infty}_{G(F(p))}(P). Using Definition 7.1 three times,

d(GF)p(X)(h)=X(h(GF))=X((hG)F)=dFp(X)(hG)=dGF(p)(dFp(X))(h).d(G\circ F)_p(X)(h) = X\big(h\circ(G\circ F)\big) = X\big((h\circ G)\circ F\big) = dF_p(X)(h\circ G) = dG_{F(p)}\big(dF_p(X)\big)(h).

As hh is arbitrary, the identity follows. For the identity map, d(id)p(X)(f)=X(fid)=X(f)d(\mathrm{id})_p(X)(f) = X(f\circ\mathrm{id}) = X(f).

Corollary 7.5The differential of a diffeomorphism is an isomorphism

If F:MNF : M \to N is a diffeomorphism, then dFp:TpMTF(p)NdF_p : T_pM \to T_{F(p)}N is a linear isomorphism at every pp, with (dFp)1=d(F1)F(p)(dF_p)^{-1} = d(F^{-1})_{F(p)}. In particular, corresponding points of diffeomorphic manifolds have tangent spaces of the same dimension.

Proof(Corollary 7.5)

Applying Theorem 7.4 to F1F=idMF^{-1}\circ F = \mathrm{id}_M and FF1=idNF\circ F^{-1} = \mathrm{id}_N gives

d(F1)F(p)dFp=idTpM,dFpd(F1)F(p)=idTF(p)N,d(F^{-1})_{F(p)}\circ dF_p = \mathrm{id}_{T_pM}, \qquad dF_p \circ d(F^{-1})_{F(p)} = \mathrm{id}_{T_{F(p)}N},

so dFpdF_p is bijective with inverse d(F1)F(p)d(F^{-1})_{F(p)}. A linear isomorphism forces equality of dimensions (see Vector spaces and linear maps, invariance of dimension(Theorem 5.5)[Vector Spaces and Linear Maps]).

Proposition 7.6The matrix of the differential is the Jacobian

Let F:MNF : M \to N be CC^{\infty}, let (U,φ)(U,\varphi) be a chart of MM containing pp (coordinates xix^i, dimM=n\dim M = n) and (V,ψ)(V,\psi) a chart of NN containing q=F(p)q = F(p) (coordinates yjy^j, dimN=m\dim N = m), and put F^=ψFφ1\widehat{F} = \psi\circ F\circ\varphi^{-1} and a=φ(p)a = \varphi(p). Then

dFp ⁣(xip)=j=1mF^jxi(a)yjq.dF_p\!\left(\left.\frac{\partial}{\partial x^i}\right|_p\right) = \sum_{j=1}^{m} \frac{\partial \widehat{F}^j}{\partial x^i}(a) \left.\frac{\partial}{\partial y^j}\right|_q .

That is, the matrix of dFpdF_p with respect to the bases {/xip}\{\partial/\partial x^i|_p\} and {/yjq}\{\partial/\partial y^j|_q\} is the Jacobian matrix DF^(a)D\widehat{F}(a) of the local representative F^\widehat{F}.

Proof(Proposition 7.6)

Take gCq(N)g \in C^{\infty}_q(N). By definition,

dFp ⁣(xip)(g)=xip(gF)=i((gF)φ1)(a).dF_p\!\left(\left.\frac{\partial}{\partial x^i}\right|_p\right)(g) = \left.\frac{\partial}{\partial x^i}\right|_p (g\circ F) = \partial_i\big((g\circ F)\circ\varphi^{-1}\big)(a).

Rewrite (gF)φ1=(gψ1)(ψFφ1)=(gψ1)F^(g\circ F)\circ\varphi^{-1} = (g\circ\psi^{-1})\circ(\psi\circ F\circ\varphi^{-1}) = (g\circ\psi^{-1})\circ\widehat{F} and apply the chain rule to the composition RnRmR\mathbb{R}^n \to \mathbb{R}^m \to \mathbb{R}:

i((gψ1)F^)(a)=j=1mj(gψ1)(F^(a))iF^j(a)=j=1miF^j(a)yjq(g)\partial_i\big((g\circ\psi^{-1})\circ\widehat F\big)(a) = \sum_{j=1}^m \partial_j (g\circ\psi^{-1})\big(\widehat F(a)\big)\cdot \partial_i \widehat F^j(a) = \sum_{j=1}^m \partial_i\widehat F^j(a)\left.\frac{\partial}{\partial y^j}\right|_q(g)

(using F^(a)=ψ(q)\widehat F(a) = \psi(q)). Since gg is arbitrary, the claim follows.

This proposition guarantees that the abstract Definition 7.1 is the same thing as the classical “linear approximation by the Jacobian matrix”. Writing Theorem 7.4 out through Proposition 7.6 returns the chain rule as a product of matrices, D(G^F^)(a)=DG^(F^(a))DF^(a)D(\widehat{G}\circ\widehat{F})(a) = D\widehat{G}(\widehat F(a))\,D\widehat F(a).

Example 7.7The tangent space of a sphere

Let Sn={uRn+1:u=1}S^n = \{u \in \mathbb{R}^{n+1} : \|u\| = 1\} be an embedded submanifold of Rn+1\mathbb{R}^{n+1} and ι:SnRn+1\iota : S^n \hookrightarrow \mathbb{R}^{n+1} the inclusion. Under the identification TpRn+1Rn+1T_p\mathbb{R}^{n+1}\cong\mathbb{R}^{n+1} of Example 6.2,

dιp(TpSn)=p={vRn+1:p,v=0}.d\iota_p(T_pS^n) = p^{\perp} = \{v \in \mathbb{R}^{n+1} : \langle p, v\rangle = 0\}.

(One inclusion.) Given [γ]TpSn[\gamma] \in T_pS^n, the curve ιγ\iota\circ\gamma lies in Rn+1\mathbb{R}^{n+1} and satisfies γ(t)2=1\|\gamma(t)\|^2 = 1. Differentiating in tt (using bilinearity of the inner product and the product rule),

0=ddtγ(t),γ(t)t=0=2γ(0),(ιγ)(0)=2p,(ιγ)(0),0 = \frac{d}{dt}\langle \gamma(t),\gamma(t)\rangle\Big|_{t=0} = 2\langle \gamma(0), (\iota\circ\gamma)'(0)\rangle = 2\langle p, (\iota\circ\gamma)'(0)\rangle,

so by Proposition 7.3 we get dιp([γ])=(ιγ)(0)pd\iota_p([\gamma]) = (\iota\circ\gamma)'(0) \in p^{\perp}.

(The other inclusion.) Let vpv \in p^{\perp} with v0v \ne 0, put e=v/ve = v/\|v\| and ω=v\omega = \|v\|, and define

γ(t)=(cosωt)p+(sinωt)e.\gamma(t) = (\cos\omega t)\,p + (\sin\omega t)\,e .

Since pp and ee are orthonormal, γ(t)2=cos2ωt+sin2ωt=1\|\gamma(t)\|^2 = \cos^2\omega t + \sin^2\omega t = 1, so γ\gamma is a curve in SnS^n; as SnS^n is an embedded submanifold, it is CC^{\infty} also as a map into SnS^n. From γ(0)=p\gamma(0) = p and γ(0)=ωe=v\gamma'(0) = \omega e = v we conclude that vv lies in the image. The vector v=0v = 0 is obtained as 0=dιp(0)0 = d\iota_p(0).

(Conclusion.) The map dιpd\iota_p is linear (Proposition 7.2), and its image both contains and is contained in pp^{\perp}, hence equals pp^{\perp}. Since dimTpSn=n=dimp\dim T_pS^n = n = \dim p^{\perp} (Theorem 5.4), a surjective linear map between spaces of equal dimension is also injective, so dιpd\iota_p is injective and TpSnpT_pS^n \cong p^{\perp}. The familiar picture of “a plane touching the sphere” is this isomorphism drawn inside Rn+1\mathbb{R}^{n+1} after translating it so that its base point is pp.

Example 7.8The tangent space of the orthogonal group at the identity

The set O(n)={AMatn(R):ATA=I}O(n) = \{A \in \mathrm{Mat}_n(\mathbb{R}) : A^{\mathsf T}A = I\} is a manifold of dimension n(n1)/2n(n-1)/2 (see Lie groups and Lie algebras, the orthogonal group O(n) and the unitary group U(n)(Example 4.4)[リー群とリー環]). Inside Matn(R)Rn2\mathrm{Mat}_n(\mathbb{R})\cong\mathbb{R}^{n^2} we have

TIO(n)={AMatn(R):A+AT=0},T_I O(n) = \{A \in \mathrm{Mat}_n(\mathbb{R}) : A + A^{\mathsf T} = 0\},

the space of skew-symmetric matrices. Indeed, a curve γ\gamma in O(n)O(n) with γ(0)=I\gamma(0) = I satisfies γ(t)Tγ(t)=I\gamma(t)^{\mathsf T}\gamma(t) = I, and differentiating at t=0t = 0 gives

γ(0)Tγ(0)+γ(0)Tγ(0)=γ(0)T+γ(0)=0.\gamma'(0)^{\mathsf T}\gamma(0) + \gamma(0)^{\mathsf T}\gamma'(0) = \gamma'(0)^{\mathsf T} + \gamma'(0) = 0.

Conversely, if AT=AA^{\mathsf T} = -A then γ(t)=exp(tA)\gamma(t) = \exp(tA) takes values in O(n)O(n), because γ(t)T=exp(tAT)=exp(tA)=γ(t)1\gamma(t)^{\mathsf T} = \exp(tA^{\mathsf T}) = \exp(-tA) = \gamma(t)^{-1}, and γ(0)=I\gamma(0) = I, γ(0)=A\gamma'(0) = A. The space of skew-symmetric matrices has dimension n(n1)/2n(n-1)/2, matching dimO(n)\dim O(n), so the same dimension count as in Example 7.7 settles the equality.

Now that we have a tangent space at each point, we bundle them together over all points. A vector field on a manifold is “a rule assigning to each point a tangent vector at that point”, but to discuss its smoothness the set of all tangent vectors must itself be a manifold.

Definition 8.1The tangent bundle

Let MM be an nn-dimensional CC^{\infty} manifold. As a set, put

TM=pMTpM={(p,X):pM, XTpM}TM = \bigsqcup_{p \in M} T_pM = \{(p, X) : p \in M,\ X \in T_pM\}

and call π:TMM\pi : TM \to M, π(p,X)=p\pi(p,X) = p, the projection. We call TMTM the tangent bundle of MM, and π1(p)=TpM\pi^{-1}(p) = T_pM the fiber over pp.

pMT_pMsection (vector field)
Schematic picture of the tangent bundle: over each point of the base M stands a fiber T_pM, and a vector field is drawn as a section threading through them.

Theorem 8.2The tangent bundle is a 2n2n-dimensional manifold

Let MM be an nn-dimensional CC^{\infty} manifold. For a CC^{\infty} atlas {(Uα,φα)}\{(U_{\alpha},\varphi_{\alpha})\} of MM, writing φα=(xα1,,xαn)\varphi_{\alpha} = (x^1_{\alpha},\ldots,x^n_{\alpha}), define

φ~α:π1(Uα)φα(Uα)×RnR2n,φ~α(p,ivixαip)=(φα(p),v).\widetilde{\varphi}_{\alpha} : \pi^{-1}(U_{\alpha}) \to \varphi_{\alpha}(U_{\alpha})\times\mathbb{R}^n \subset \mathbb{R}^{2n}, \qquad \widetilde{\varphi}_{\alpha}\Big(p, \sum_i v^i \left.\frac{\partial}{\partial x^i_{\alpha}}\right|_p\Big) = \big(\varphi_{\alpha}(p), v\big).

Then there is exactly one topology and CC^{\infty} structure on TMTM having {(π1(Uα),φ~α)}\{(\pi^{-1}(U_{\alpha}), \widetilde\varphi_{\alpha})\} as an atlas, and TMTM becomes a CC^{\infty} manifold of dimension 2n2n. Moreover π:TMM\pi : TM \to M is a surjective CC^{\infty} submersion.

Proof(Theorem 8.2)

(1) Each φ~α\widetilde\varphi_{\alpha} is a bijection. By Theorem 5.4, every tangent vector at pUαp \in U_{\alpha} has a unique component expression in the basis {/xαip}\{\partial/\partial x^i_{\alpha}|_p\}, so π1(Uα)φα(Uα)×Rn\pi^{-1}(U_{\alpha}) \to \varphi_{\alpha}(U_{\alpha})\times\mathbb{R}^n is bijective.

(2) Transition maps. Let Uαβ=UαUβU_{\alpha\beta} = U_{\alpha}\cap U_{\beta} \ne \emptyset and put τ=φβφα1\tau = \varphi_{\beta}\circ\varphi_{\alpha}^{-1}. The set φ~α(π1(Uαβ))=φα(Uαβ)×Rn\widetilde\varphi_{\alpha}(\pi^{-1}(U_{\alpha\beta})) = \varphi_{\alpha}(U_{\alpha\beta})\times\mathbb{R}^n is open in R2n\mathbb{R}^{2n}. By the change-of-basis formula established in the proof of Theorem 6.1,

φ~βφ~α1(x,v)=(τ(x), Dτ(x)v).\widetilde\varphi_{\beta}\circ\widetilde\varphi_{\alpha}^{-1}(x, v) = \big(\tau(x),\ D\tau(x)\,v\big).

Since τ\tau is CC^{\infty}, every entry of xDτ(x)x \mapsto D\tau(x) is CC^{\infty}, and Dτ(x)vD\tau(x)v is CC^{\infty} in xx and linear in vv, hence CC^{\infty} altogether. The inverse map has the same form with α\alpha and β\beta interchanged, so the transition maps are diffeomorphisms.

(3) Topology. We check that B={φ~α1(W):α, Wφα(Uα)×Rn open}\mathcal{B} = \{\widetilde\varphi_{\alpha}^{-1}(W) : \alpha,\ W \subset \varphi_{\alpha}(U_{\alpha})\times\mathbb{R}^n \text{ open}\} is a basis for a topology on TMTM. It covers TMTM. Take a point ξ\xi in the intersection of two members φ~α1(W1)\widetilde\varphi_{\alpha}^{-1}(W_1) and φ~β1(W2)\widetilde\varphi_{\beta}^{-1}(W_2); then ξπ1(Uαβ)\xi \in \pi^{-1}(U_{\alpha\beta}), and by (2) the set φ~α(φ~α1(W1)φ~β1(W2))=W1(φ~βφ~α1)1(W2)\widetilde\varphi_{\alpha}\big(\widetilde\varphi_{\alpha}^{-1}(W_1)\cap\widetilde\varphi_{\beta}^{-1}(W_2)\big) = W_1 \cap (\widetilde\varphi_{\beta}\circ\widetilde\varphi_{\alpha}^{-1})^{-1}(W_2) is open; calling it W3W_3 we get ξφ~α1(W3)\xi \in \widetilde\varphi_{\alpha}^{-1}(W_3) \subset the intersection. So B\mathcal{B} is a basis, and in the topology it generates each φ~α\widetilde\varphi_{\alpha} is a homeomorphism.

(4) Hausdorff. Let ξη\xi \ne \eta in TMTM. If π(ξ)=π(η)=p\pi(\xi) = \pi(\eta) = p, choose α\alpha with pUαp \in U_{\alpha}; then φ~α(ξ)φ~α(η)\widetilde\varphi_{\alpha}(\xi)\ne\widetilde\varphi_{\alpha}(\eta), and these can be separated inside φα(Uα)×Rn\varphi_\alpha(U_\alpha)\times\mathbb{R}^n because R2n\mathbb{R}^{2n} is Hausdorff; the preimages are the required open sets. If π(ξ)π(η)\pi(\xi)\ne\pi(\eta), then since MM is Hausdorff there are disjoint open sets Uπ(ξ)U \ni \pi(\xi) and Vπ(η)V \ni \pi(\eta), and π1(U)\pi^{-1}(U) and π1(V)\pi^{-1}(V) are disjoint open sets (π\pi is continuous in the topology of (3) because its local representative is the continuous map (x,v)x(x,v)\mapsto x).

(5) Second countability. As MM is second countable, there is a countable subatlas {(Uαk,φαk)}kN\{(U_{\alpha_k},\varphi_{\alpha_k})\}_{k\in\mathbb{N}}. Each π1(Uαk)\pi^{-1}(U_{\alpha_k}) is homeomorphic to an open subset of R2n\mathbb{R}^{2n} and hence second countable, and a space covered by countably many second countable open subspaces is second countable.

(6) The CC^{\infty} structure and π\pi. By (2), {(π1(Uα),φ~α)}\{(\pi^{-1}(U_{\alpha}),\widetilde\varphi_{\alpha})\} is a CC^{\infty} atlas, and the maximal atlas containing it determines the CC^{\infty} structure. Both the topology and the atlas are forced by the requirement that each φ~α\widetilde\varphi_\alpha be a diffeomorphism, so they are unique. The local representative of π\pi is the projection φαπφ~α1(x,v)=x\varphi_{\alpha}\circ\pi\circ\widetilde\varphi_{\alpha}^{-1}(x,v) = x, which is CC^{\infty} with surjective differential (its Jacobian matrix is (In0)\begin{pmatrix} I_n & 0\end{pmatrix}), so π\pi is a submersion. Surjectivity follows because every TpMT_pM contains the zero vector.

Remark 8.3Vector fields as sections

A CC^{\infty} map X:MTM\mathcal{X} : M \to TM with πX=idM\pi\circ\mathcal{X} = \mathrm{id}_M is called a section of TMTM, and this is the official definition of a vector field. On a chart we may write X=iXi/xi\mathcal{X} = \sum_i \mathcal{X}^i \partial/\partial x^i, and, read in the charts of Theorem 8.2, smoothness of X\mathcal{X} is equivalent to smoothness of the components Xi\mathcal{X}^i. Vector fields, and the differential forms dual to them, are treated in Vector fields and differential forms (the official definition as sections is sections of a bundle and vector fields(Definition 3.1)[ベクトル場と微分形式], the dual side is cotangent spaces, the cotangent bundle and 1-forms(Definition 4.1)[ベクトル場と微分形式]).

Example 8.4TS1TS^1 is trivial

Let S1R2S^1 \subset \mathbb{R}^2 and identify TpS1pT_pS^1 \cong p^{\perp} using Example 7.7, where p=(p1,p2)p = (p_1,p_2). The line pp^{\perp} is one-dimensional, with basis Jp:=(p2,p1)Jp := (-p_2, p_1) (indeed p,Jp=p1p2+p2p1=0\langle p, Jp\rangle = -p_1p_2 + p_2p_1 = 0 and Jp=10\|Jp\| = 1 \ne 0). So define

Ψ:S1×RTS1,Ψ(p,t)=(p, tJp).\Psi : S^1\times\mathbb{R} \to TS^1, \qquad \Psi(p,t) = \big(p,\ t\,Jp\big).

Then Ψ\Psi is a bijection, since ttJpt \mapsto tJp is an isomorphism on each fiber. For smoothness, take the chart θ(cosθ,sinθ)\theta \mapsto (\cos\theta,\sin\theta) of S1S^1: the local representative of Ψ\Psi becomes (θ,t)(θ,t)(\theta,t)\mapsto(\theta,t). Indeed, in these coordinates /θp=(sinθ,cosθ)=Jp\partial/\partial\theta|_p = (-\sin\theta,\cos\theta) = Jp, so the component of tJptJp is exactly tt. The inverse is CC^{\infty} for the same reason, so Ψ\Psi is a diffeomorphism and TS1S1×RTS^1 \cong S^1\times\mathbb{R}. Equivalently, S1S^1 carries a nowhere vanishing vector field, namely pJpp \mapsto Jp.

Remark 8.5Nontrivial tangent bundles

What happens in Example 8.4 is not typical of all manifolds. A manifold with TMM×RnTM \cong M\times\mathbb{R}^n is called parallelizable, and S2S^2 is not parallelizable. This is a consequence of the hairy ball theorem (Poincaré–Brouwer): every continuous vector field on S2S^2 vanishes somewhere. Among spheres, only S1S^1, S3S^3 and S7S^7 are known to be parallelizable (Bott–Milnor and Kervaire, 1958). That the question of triviality of the tangent bundle contains deep topology is one of the reasons for introducing TMTM at all.

Exercise 9.1Easy

On M=R2M = \mathbb{R}^2, besides the usual coordinates (x,y)(x,y), introduce the global chart ψ=(u,v)\psi = (u,v) given by u=x+yu = x+y and v=xyv = x-y. At a point pp, express /up\partial/\partial u|_p and /vp\partial/\partial v|_p in terms of /xp\partial/\partial x|_p and /yp\partial/\partial y|_p. Then check the expression on the function f(x,y)=xyf(x,y) = xy.

Solution

Since ψ1(u,v)=(u+v2,uv2)\psi^{-1}(u,v) = \left(\frac{u+v}{2}, \frac{u-v}{2}\right), the Jacobian matrix of τ=φψ1\tau = \varphi\circ\psi^{-1} (with φ\varphi the Cartesian chart) is

Dτ=(1/21/21/21/2).D\tau = \begin{pmatrix} 1/2 & 1/2 \\ 1/2 & -1/2 \end{pmatrix}.

By the transformation formula in the proof of Theorem 6.1 (namely /ui=jiτj/xj\partial/\partial u^i = \sum_j \partial_i\tau^j \cdot \partial/\partial x^j),

up=12(xp+yp),vp=12(xpyp).\left.\frac{\partial}{\partial u}\right|_p = \frac12\left(\left.\frac{\partial}{\partial x}\right|_p + \left.\frac{\partial}{\partial y}\right|_p\right), \qquad \left.\frac{\partial}{\partial v}\right|_p = \frac12\left(\left.\frac{\partial}{\partial x}\right|_p - \left.\frac{\partial}{\partial y}\right|_p\right).

Now the check. In the new coordinates f(x,y)=xyf(x,y) = xy becomes fψ1(u,v)=u+v2uv2=u2v24f\circ\psi^{-1}(u,v) = \frac{u+v}{2}\cdot\frac{u-v}{2} = \frac{u^2-v^2}{4}, so /up(f)=u/2\partial/\partial u|_p(f) = u/2 and /vp(f)=v/2\partial/\partial v|_p(f) = -v/2. Computing with the right-hand sides gives 12(y+x)=x+y2=u2\frac12(y + x) = \frac{x+y}{2} = \frac{u}{2} and 12(yx)=xy2=v2\frac12(y - x) = -\frac{x-y}{2} = -\frac v2, in agreement.

Exercise 9.2Standard

Let XTpMX \in T_pM be a derivation and let f,gCp(M)f, g \in C^{\infty}_p(M) satisfy f(p)=g(p)=0f(p) = g(p) = 0. Show that X(fg)=0X(fg) = 0. Then, setting mp={fCp(M):f(p)=0}\mathfrak{m}_p = \{f \in C^{\infty}_p(M) : f(p) = 0\} and mp2={k=1Nfkgk:fk,gkmp}\mathfrak{m}_p^2 = \{\sum_{k=1}^N f_kg_k : f_k,g_k\in\mathfrak{m}_p\}, show that XXmpX \mapsto X|_{\mathfrak{m}_p} gives a linear isomorphism TpM(mp/mp2)T_pM \to (\mathfrak{m}_p/\mathfrak{m}_p^2)^{*}.

Solution

The first claim is an immediate consequence of the Leibniz rule: X(fg)=X(f)g(p)+f(p)X(g)=X(f)0+0X(g)=0X(fg) = X(f)g(p) + f(p)X(g) = X(f)\cdot 0 + 0\cdot X(g) = 0. By linearity, XX vanishes on every element of mp2\mathfrak{m}_p^2, these being finite sums.

Now the isomorphism. Restricting XTpMX \in T_pM to mp\mathfrak{m}_p gives, by the above, a map vanishing on mp2\mathfrak{m}_p^2, hence an induced linear functional Xˉ\bar X on the quotient mp/mp2\mathfrak{m}_p/\mathfrak{m}_p^2. The correspondence XXˉX \mapsto \bar X is clearly linear, restriction and passage to a quotient both being linear operations.

Injectivity. Suppose Xˉ=0\bar X = 0, i.e. Xmp=0X|_{\mathfrak{m}_p} = 0. Any fCp(M)f \in C^{\infty}_p(M) decomposes as f=f(p)+(ff(p))f = \underline{f(p)} + (f - \underline{f(p)}), with the second term in mp\mathfrak{m}_p. By Lemma 5.2 we have X(f(p))=0X(\underline{f(p)}) = 0, so X(f)=X(ff(p))=0X(f) = X(f - \underline{f(p)}) = 0. Hence X=0X = 0.

Surjectivity. Given λ(mp/mp2)\lambda \in (\mathfrak{m}_p/\mathfrak{m}_p^2)^{*}, define X(f):=λ([ff(p)])X(f) := \lambda\big([f - \underline{f(p)}]\big). Linearity is clear. For the Leibniz rule, decompose fgf(p)g(p)=f(p)(gg(p))+g(p)(ff(p))+(ff(p))(gg(p))fg - \underline{f(p)g(p)} = f(p)\,(g - \underline{g(p)}) + g(p)\,(f - \underline{f(p)}) + (f-\underline{f(p)})(g-\underline{g(p)}) and note that the last term lies in mp2\mathfrak{m}_p^2, hence is killed by λ\lambda. That this XX induces λ\lambda is clear from the definition.

Combined with Theorem 5.4, this shows dimmp/mp2=n\dim \mathfrak{m}_p/\mathfrak{m}_p^2 = n, and mp/mp2\mathfrak{m}_p/\mathfrak{m}_p^2 is identified with the cotangent space TpMT_p^{*}M. From this angle, Lemma 5.3 says that every element of mp\mathfrak{m}_p is a linear combination of the coordinate functions plus an element of mp2\mathfrak{m}_p^2.

Exercise 9.3Standard

The set GL(n,R)={AMatn(R):detA0}\mathrm{GL}(n,\mathbb{R}) = \{A \in \mathrm{Mat}_n(\mathbb{R}) : \det A \ne 0\} is open in Matn(R)Rn2\mathrm{Mat}_n(\mathbb{R})\cong\mathbb{R}^{n^2} and hence a manifold of dimension n2n^2. Under the identification TIGL(n,R)Matn(R)T_I\mathrm{GL}(n,\mathbb{R})\cong\mathrm{Mat}_n(\mathbb{R}), show that the differential of det:GL(n,R)R\det : \mathrm{GL}(n,\mathbb{R})\to\mathbb{R} at the identity is d(det)I(A)=trAd(\det)_I(A) = \operatorname{tr}A.

Solution

Since GL(n,R)\mathrm{GL}(n,\mathbb{R}) is open in Rn2\mathbb{R}^{n^2}, taking the inclusion as a chart gives TIGL(n,R)Matn(R)T_I\mathrm{GL}(n,\mathbb{R}) \cong \mathrm{Mat}_n(\mathbb{R}) (the same argument as in Example 6.2). For AMatn(R)A \in \mathrm{Mat}_n(\mathbb{R}) take the curve γ(t)=I+tA\gamma(t) = I + tA. As det\det is continuous, detγ(t)0\det\gamma(t)\ne 0 for t|t| small enough, so γ\gamma is a curve in GL(n,R)\mathrm{GL}(n,\mathbb{R}) with γ(0)=I\gamma(0) = I and γ(0)=A\gamma'(0) = A. By Proposition 7.3, d(det)I(A)=(detγ)(0)d(\det)_I(A) = (\det\circ\gamma)'(0).

Expand det(I+tA)\det(I+tA). By the definition of the determinant (see Determinants, the Leibniz formula(Definition 3.1)[Determinants and Their Properties]),

det(I+tA)=σSnsgn(σ)i=1n(δiσ(i)+tAiσ(i)).\det(I+tA) = \sum_{\sigma\in S_n}\operatorname{sgn}(\sigma)\prod_{i=1}^{n}\big(\delta_{i\sigma(i)} + tA_{i\sigma(i)}\big).

Collect the terms of degree at most 11 in tt. If σ\sigma is not the identity permutation, then σ(i)i\sigma(i)\ne i for at least two values of ii (a permutation moves at least two points), so the corresponding product is divisible by t2t^2. The identity permutation contributes i(1+tAii)=1+tiAii+O(t2)\prod_i (1 + tA_{ii}) = 1 + t\sum_i A_{ii} + O(t^2). Hence

det(I+tA)=1+ttrA+O(t2),\det(I+tA) = 1 + t\operatorname{tr}A + O(t^2),

and differentiating in tt at t=0t=0 gives d(det)I(A)=trAd(\det)_I(A) = \operatorname{tr}A.

Exercise 9.4Hard

Let MM be a connected CC^{\infty} manifold, NN a CC^{\infty} manifold and F:MNF : M \to N a CC^{\infty} map. Show that if dFp=0dF_p = 0 for every pMp \in M, then FF is constant.

Solution

Fix p0Mp_0\in M arbitrarily, put q0=F(p0)q_0 = F(p_0), and show that S=F1(q0)S = F^{-1}(q_0) is nonempty, open and closed. Connectedness of MM (characterization of connectedness(Proposition 3.1)[連結性]: a nonempty subset that is both open and closed must be the whole space) then gives S=MS = M, i.e. FF is constant.

SS \ne \emptyset because p0Sp_0 \in S. And SS is closed because {q0}\{q_0\} is closed (NN being Hausdorff) and FF is continuous.

Now we show SS is open. Take pSp \in S and choose a chart (U,φ)(U,\varphi) containing pp and a chart (V,ψ)(V,\psi) containing F(p)=q0F(p) = q_0 such that F(U)VF(U)\subset V and φ(U)\varphi(U) is an open ball, hence connected (by continuity of FF it suffices to shrink UU to UF1(V)U\cap F^{-1}(V)). Put F^=ψFφ1:φ(U)ψ(V)\widehat F = \psi\circ F\circ\varphi^{-1} : \varphi(U)\to\psi(V). By Proposition 7.6, the condition dFx=0dF_x = 0 for xUx\in U is equivalent to the vanishing of the Jacobian matrix DF^(φ(x))D\widehat F(\varphi(x)). By hypothesis this holds on all of φ(U)\varphi(U), so all partial derivatives of every component of F^\widehat F vanish. Since φ(U)\varphi(U) is a connected open set, F^\widehat F is constant (componentwise, either join points by segments and apply the mean value theorem, or quote it as a corollary in Differentiation in several variables). As F^(φ(p))=ψ(q0)\widehat F(\varphi(p)) = \psi(q_0), we get F^ψ(q0)\widehat F \equiv \psi(q_0), that is FUq0F|_U \equiv q_0, so USU \subset S. Hence SS is open.

Dropping the hypothesis that MM is connected destroys the conclusion. For instance take M=RRM = \mathbb{R}\sqcup\mathbb{R} (the disjoint union of two lines) and N=RN=\mathbb{R}, and let FF be 00 on one copy and 11 on the other: then dFp=0dF_p = 0 at every point, yet FF is not constant.

  • Matsumoto Yukio, Tayōtai no Kiso, University of Tokyo Press, 1988 (in Japanese) — an elementary introduction to tangent vectors and tangent spaces. A representative example of the exposition that starts from the definition by curves.
  • J. M. Lee, Introduction to Smooth Manifolds, 2nd ed., Springer GTM 218, 2013 — Chapter 3 (Tangent Vectors). The definition by derivations, the differential, and the construction of the tangent bundle are developed in almost the same order as in this article.
  • F. W. Warner, Foundations of Differentiable Manifolds and Lie Groups, Springer GTM 94, 1983 — Chapter 1. A concise account of the definition by derivations and of why smoothness is needed.
  • M. Spivak, A Comprehensive Introduction to Differential Geometry, Vol. 1, 3rd ed., Publish or Perish, 1999 — Chapter 3. A detailed side-by-side comparison of the three definitions, with the role of Hadamard’s lemma made explicit.
  • R. Bott and J. Milnor, “On the parallelizability of the spheres”, Bulletin of the American Mathematical Society 64 (1958), 87–89 — the result on parallelizability of spheres mentioned in Remark 8.5.

Appendix: Why the space of derivations grows in the CkC^k setting

Section titled “Appendix: Why the space of derivations grows in the CkC^kCk setting”

Where the problem lies. What was essential in the proof of Theorem 5.4 was Lemma 5.3, that is, the fact that an element of mp\mathfrak{m}_p falls into mp2\mathfrak{m}_p^2 once linear combinations of the coordinate functions are removed. As we saw in Exercise 9.2, the space of derivations at pp is identified with (mp/mp2)(\mathfrak{m}_p/\mathfrak{m}_p^2)^{*}. So the statement “the space of derivations is nn-dimensional” and the statement ”mp/mp2\mathfrak{m}_p/\mathfrak{m}_p^2 is nn-dimensional” are one and the same.

What goes wrong for C1C^1. Take n=1n=1 and p=0p=0, and consider the algebra C01(R)C^1_0(\mathbb{R}) of germs of C1C^1 functions. Let m\mathfrak{m} be the germs vanishing at 00 and m2\mathfrak{m}^2 the finite sums of products of such. For 0<α<10 < \alpha < 1 consider fα(x)=x1+αf_{\alpha}(x) = |x|^{1+\alpha}. It is C1C^1 (its derivative fα(x)=(1+α)sgn(x)xαf_{\alpha}'(x) = (1+\alpha)\operatorname{sgn}(x)|x|^{\alpha} is continuous at 00 with value 00) and lies in m\mathfrak{m}. On the other hand, if g,hmg,h\in\mathfrak{m} are C1C^1, then the mean value theorem gives g(x)Cx|g(x)| \le C|x| and h(x)Cx|h(x)|\le C|x| near 00, so any finite sum kgkhk\sum_k g_kh_k is O(x2)O(x^2). But x1+α/x2|x|^{1+\alpha}/x^2 \to \infty as x0x\to 0 when α<1\alpha<1, so fαm2f_{\alpha}\notin\mathfrak{m}^2. Moreover, the same comparison of growth rates shows that the fαf_{\alpha} for distinct α\alpha are linearly independent in m/m2\mathfrak{m}/\mathfrak{m}^2. Hence m/m2\mathfrak{m}/\mathfrak{m}^2 is infinite-dimensional, and so is its dual, the space of derivations.

Conclusion. In the CC^{\infty} world a small miracle occurs: mp/mp2\mathfrak{m}_p/\mathfrak{m}_p^2 collapses to exactly nn dimensions, and what guarantees this is Lemma 5.3. When working with CkC^k manifolds for finite kk, the standard prescription is to adopt Definition 3.2 or Definition 4.1 as the definition of the tangent space and to give up the characterization by derivations. A detailed discussion of this state of affairs can be found in Chapter 1 of Warner.

Report an error in this article ・Operated by: Mugen Giken LLCPricingTermsLegal notice

© 2026 夢現技研合同会社 ・Feeding the text to an LLM is welcome. Code samples are MIT licensed.