Differential calculus is essential to the study of differential geometry - that's the differential part!
Here we review some of the most relevant topics. Some topics may not be covered in standard courses, and the perspective taken as well as some of the notation may also be unfamiliar. Thus even if you have a solid background in the material, it's worth at least taking a look. Throughout the course if you encounter something you're not comfortable with, you can always return here to see if it's explained here.
Let \(U \subseteq \mathbb{R}^n\) be an open set and let \(f : U \to \mathbb{R}^k\) be a function. We say that \(f\) has a directional derivative \(D_X f(p)\) in direction \(X\) a the point \(p\) if the limit \[ \lim_{t \to 0} \frac{f(p + tX) - f(p)}{t} \] exists. In that case we write \(D_X f(p)\) for the limit.
Let \(U \subseteq \mathbb{R}^n\) be an open set and let \(f = (f^1, \dots, f^k) : U \to \mathbb{R}^k\). Then \(D_X f(p)\) exists if and only if for each \(a = 1, \dots, k\), \(D_x f^a\) exists. In this case, \[ D_X f(p) = \big(D_X f^1 (p), \dots, D_X f^k(p)\big). \]
Observe that \(D_X f (p)\) exists if and only if \[ \lim_{t\to 0} \left\|\frac{f(p + tX) - f(p)}{t} - D_X f(p)\right\| = 0. \]
On the other hand \(D_X f^a (p)\) exists if and only if \[ \lim_{t\to 0} \left\|\frac{f^a(p + tX) - f^a(p)}{t} - D_X f^a(p)\right\| = 0. \] where \((D_X f(p))^i = \pi_i (D_X f(p))\) is the \(i\)'th component of \(D_X f(p)\).
That \([D_X f(p)]^a = D_X f^a (p)\) follows since limits are unique.
Fill in the details of the proof.
Let \(f(x, y) = (x^2 y, x - y)\). Then for \(p = (x, y)\) and \(X = (u, v)\),
\begin{align*} \frac{f(p + tX) - f(p)}{1} &= \frac{1}{t} \big((x + tu)^2 (y+tv) - x^2 y, x + tu - y - tv - x - y\big) \\ &= \frac{1}{t}\big(t x^2v + 2t xyu + 2t^2xuv + t^2yu + t^3 uv, tu - tv\big). \end{align*}Taking the limit \(t\to 0\) we obtain \[ D_X f(p) = \big(x^2 v + 2xyu, u - v\big). \]
Show that \(D_X f(p)\) exists if and only if \(\partial_t|_{t=0} f^a (p + tX)\) exists for each \(a = 1, \dots, k\). In that case \(D_X f(p) = \big((\pi^1 \circ f)'(0), \dots, (\pi^k \circ f)'(0)\big)\). Here \(\pi^a \circ f = f^a\) is the \(a\)'th component of \(f\).
Thus we may compute \(D_X f (p)\) be computing derivatives of the scalar valued functions, \(\pi^a \circ f(p + tX)\) of the single variable \(t\).
Let \(U \subseteq \mathbb{R}^n\) be an open set and let \(p \in U\). Let \(f : U \to \mathbb{R}^k\) be a function such that \(D_X f (p)\) exists for every \(X \in \mathbb{R}^n\) and such that the map \(X \mapsto D_X f(p)\) is linear in \(X\). The differential , \(d_p f\) is the map \(\mathbb{R}^n \to \mathbb{R}^k\) defined by \[ df_p (X) = D_X f (p). \]
We say that \(f\) has all directional derivatives at \(p\) with differential \(df_p\). Sometimes we say that \(f\) is Gateaux differentiable at \(p\).
Important Remark! There are some subtleties here. It's possible for example, that for some functions, \(df_p\) is not a linear map! We don't want that, so we restrict to functions \(f\) where \(df_p\) is linear. We won't go into these subtleties here. For the interested reader, the derivative we've defined is the so-called Gateaux derivative .
Even with the assumption of linearity, we will make a further restriction to avoid some other undesirable behaviours (again we won't go into those here).
Let \(U \subseteq \mathbb{R}^n\) be an open set and let \(f : U \to \mathbb{R}^k\) be a function such that for every \(p \in U\), \(f\) has all directional derivatives (so that \(df_p\) is defined). We say that \(f\) is \(C^1\) provided the map \(f \mapsto df_p\) is continuous.
There are a few equivalent ways to interpret continuity here. Here's one way: let \(\{e_i\}_{i=1}^n\) be a basis for \(\mathbb{R}^n\) and let \(\{u_a\}_{a=1}^k\) be a basis for \(\mathbb{R}^k\). Then with respect to these bases, \(df_p\) is the \(k \times n\) matrix \[ [df_p]_i^a = \theta^a (df_p(e_i)) \] where \(\{\theta^a\}\) is the dual basis to \(\{u_a\}\). That is the \(i,a\) entry of the matrix is the \(a\)'th component of \(df_p(e_i)\). Then \(p \mapsto [df_p]_i^a\) is a matrix valued map, which we interested as a map from \(U \subseteq \mathbb{R}^n\) to \(\mathbb{R}^{nk}\). We require this map to be a continuous map between Euclidean spaces.
Let \(f(x, y) = (x, e^{y - x}, \sin(xy))\). Let \(e_1 = (1, 0)\), \(e_2 = (0, 1)\) be the standard bases for \(\mathbb{R}^2\) and similarly let \(u_1, u_2, u_3\) be the standard basis for \(\mathbb{R}^3\). Then for \(p = (x, y)\),
\begin{align*} df_p (e_1) &= \partial_t|_{t=0} f(p + te_1) \\ &= \partial_t|_{t=0} f(x + t, y) \\ &= \partial_t|_{t=0} (x+t, e^{y-x-t}, \sin((x+t)y) \\ &= (1, -e^{-y-x}, y\cos(x)) \\ &= u_1 - e^{y-x} u_2 + y\cos(xy) u_3. \end{align*}Similarly,
\begin{align*} df_p (e_2) &= \partial_t|_{t=0} f(p + te_2) \\ &= \partial_t|_{t=0} f(x, y + t) \\ &= \partial_t|_{t=0} (x, e^{y+t-x}, \sin(x(y+t)) \\ &= (0, e^{y-x}, x\cos(xy)) \\ &= e^{y-x} u_2 + x \cos(xy) u_3. \end{align*}Thus the differential at \(p = (x, y)\) is the linear map
\begin{align*} df_p(X^1 e_1 + X^2 e_2) &= X^1 df_p(e_1) + X^2 df_p(e_2) \\ &= X^1 \left(u_1 - e^{y-x} u_2 + y\cos(xy) u_3\right) + X^2 \left(e^{y-x} u_2 + x \cos(xy) u_3\right) \\ &= X^1 u_1 + e^{y-x} (X^2 - X^1) u_2 + \cos(xy)(yX^1 + x X^2) u_3. \end{align*}In matrix form, \((df_p)_i^a = \theta^a(df_p(e_i))\) is \[ df_p = \begin{pmatrix} 1 & 0 \\ -e^{y-x} & e^{y-x} \\ y \cos(xy) & x \cos(xy) \end{pmatrix}. \]
Let \(f : \mathbb{R}^n \to \mathbb{R}^k\) be a \(C^1\) map. The \(i\)'th partial derivative, \(\partial_i f\) is defined to be \[ \partial_i f(p) = df_p(e_i). \]
Let \(f = (f^1, \dots, f^m) : \mathbb{R}^n \to \mathbb{R}^k\) be a \(C^1\) map. Let \(\{e_1, \dots, e_n\}\) and \(\{u_1, \dots, u_k\}\) be the standard bases. Then at \(p \in \mathbb{R}^n\), \[ (df_p)^i_a := \theta^a (df_p(e_i)) = \partial_i f^a. \] In matrix form \[ df_p = \begin{pmatrix} \partial_1 f^1 (p) & \dots & \partial_n f^1 (p) \\ \vdots & \ddots & \vdots \\ \partial_1 f^k (p) & \dots & \partial_n f^k (p) \end{pmatrix}. \]
In particular, \[ df_p = df^1_p u_1 + \dots + df^k_p u_k. \]
This follows since for a linear transformation \(T : V \to W\), with respect to bases for \(V\) and \(W\), \(T\) is represented as the matrix \(\theta^a(T(e_i))\).
If the proof is not clear, work through the details!
The rows of \(df_p\) are \((\partial_1 f^a, \dots, \partial_n f^a)\), \(a = 1, \dots, k\) while the columns are \((\partial_i f^1, \dots, \partial_i f^m)\), \(i = 1, \dots, n\).
There is another approach to multi-variable differentiation know as Frechet Differentiation . First we need to quantify one function vanishing faster than another.
Let \(f : U \subseteq \mathbb{R}^n \to \mathbb{R}^k\) be a function with \(U\) open and let \(p_0 \in U\). Let \(g : V \backslash \{p_0\} \to \mathbb{R}\) be a strictly positive function where \(V \subseteq U\) is open with \(p_0 \in V\). We write \[ f = o(g) \text{ as } p \to p_0 \] if \[ \lim_{p \to p_0} \frac{f(p)}{g(p)} = 0. \]
The idea is that \(f = o(g)\) means that \(f\) goes to \(0\) faster than \(g\) does as \(p \to p_0\).
Let \(f(x) = x^2 - 3x^5\). Then \(f = o(1)\) and \(f = o(x)\) as \(x \to 0\). The former is since \(\lim_{x\to 0} f(x) = 0\) while the latter is since \(f(x)/x = x - 3x^4 \to 0\) as \(x \to 0\). Other the other hand \(f\) is not \(o(x^2)\) since \(f(x)/x^2 \to 1\) as \(x \to 0\).
Let \(f : U \subseteq \mathbb{R}^n \to \mathbb{R}^k\) be a function with \(U\) open. We say that \(f\) is differentiable at \(p \in U\) if there is a linear map \(L_p : \mathbb{R}^n \to \mathbb{R}^k\) such that \[ f(q) = f(p) + L_p (q-p) + o(\|q-p\|) \] as \(q \to p\).
If \(f\) is differentiable at \(p\), we write \(df_p\) for the linear map \(L_p\). See the theorem below.
Here the function \(q \mapsto f(p) + L_p (q-p)\) is the linear approximation of \(f\) at \(p\). Letting \(R_p(q) = f(q) - [f(p) + L_p (q-p)]\) be the remainder (or error) after approximation, then the statement that \(f\) is differentiable at \(p\) is precisely the statement that the remainder \(R_p(q-p) = o(\|q-p\|)\). In other words, \(f\) is differentiable at \(p\) precisely when it may be approximated by a degree one polynomial with error vanishing faster than linear.
If \(f\) is differentiable at \(p\), then \(L_p\) is unique.
Suppose \[ R(q) := f(q) - [f(p) + L (q-p)] = o(\|q-p\|) \] and \[ S(q) := f(q) - [f(p) + K (q-p)] = o(\|q-p\|). \] for linear maps \(L, K\). Then \[ \frac{(R-S)(q)}{\|q-p\|} = \frac{R(q)}{\|q-p\|} - \frac{S(q)}{\|q-p\|} \to 0 \] as \(q \to p\), hence \(R - S = o(\|q-p\|)\). Since \[ (R - S) (q-p) = L (q - p) - K (q - p) = (L - K)(q - p), \] we have that
\begin{align*} (L - K) \left(\frac{q-p}{\|q-p\|}\right) &= \frac{(L - K)(q-p)}{\|q-p\|} \\ &= \frac{(R - S)(q-p)}{\|q-p\|} \to 0 \end{align*}as \(q \to p\).
Now suppose it was the case that \(L \neq K\). Then there would be a \(V \in \mathbb{R}^n\) such that \((K-L)(V) \neq 0\). Then \[ \lim_{t \to 0^+} (L - K) \left(\frac{tV}{\|tV\|}\right) = \frac{1}{\|V\|} (L - K)(V) \neq 0. \] But now letting \(q(t) = p + tV\) we get \(\lim_{t \to 0} q(t) = p\) and \[ \lim_{t\to 0^+} (L - K) \left(\frac{q-p}{\|q-p\|}\right) = \lim_{t \to 0^+} (L - K) \left(\frac{tV}{\|tV\|}\right) \neq 0. \] This contradicts that \(\lim_{q \to p} (L - K)(q - p) = 0\), hence \(L = K\) proving uniqueness.
If \(f\) is differentiable at \(p\), then \(f\) has all directional derivatives at \(p\). In this case, \(df_p = L_p\).
Here \(df_p\) denotes the differential defined via the directional derivative, \(df_p(V) = D_V f (p)\).
For any \(V \in \mathbb{R}^n\) with \(\|V\| = 1\), letting \(q(t) = p + tV\) we have \(q(t) - p = tV\), and \(\|q(t) - p\| = |t|\). Hence
\begin{align*} \left\|\frac{f(p) - f(p + t V)}{t} - L_p(V)\right\| &= \left\|\frac{f(p) - f(p + t V) - L_p(tV)}{t}\right\| \\ &= \frac{\|f(p) - f(q(t)) - L_p(q(t) - p)\|}{\|q(t) - p\|}. \end{align*}Since \(f\) is differentiable at \(p\), and \(\lim_{t\to 0} q(t) = q\), the right hand side goes to \(0\) as \(t \to 0\). Thus \[ \lim_{t \to 0} \frac{f(p) - f(p + t V)}{t} = L_p(V). \] But the left hand side is precisely \(D_V f(p) = df_p (V)\). Thus for any unit length vector \(V\), \(L_p(V) = df_p(V)\). For any \(V \neq 0\) (not necessarily unit length) by linearity we then have \[ df_p(V) = \|V\| df_p\left(\frac{V}{\|V\|}\right) = \|V\| L_p\left(\frac{V}{\|V\|}\right) = L_p(V). \] For \(V = 0\) we have \(df_p(0) = 0 = L_p(V)\) by linearity.
Thus for any \(V\), \(df_p(V) = L_p(V)\) hence \(df_p = L_p\).
Important Warning! If \(f\) has all directional derivatives, it may fail to be differentiable! A further condition is needed to ensure \(f\) is differentiable.
Let \(f\) be \(C^1\). Then \(f\) is differentiable at each point in its domain.
The proof is based on the mean value theorem. We omit it here.
Let \(f : U \subseteq \mathbb{R}^m \to \mathbb{R}^k\) and \(g : V \subseteq \mathbb{R}^n \to \mathbb{R}^m\) be functions with \(U, V\) open. Let \(p \in g^{\ast} (U)\) (i.e. \(g(p) \in U\)). If \(g\) is differentiable at \(p\) and \(f\) is differentiable at \(g(p)\), then \(f \circ g\) is differentiable at \(p\). In this case, \[ d(f \circ g)_p = df_{g(p)} \circ dg_p. \]
We will show that \(f \circ g\) is differentiable at \(p\) by showing that \(d(f \circ g)_p = df_{g(p)} \circ dg_p\). That is, we aim to show for any \(V \in \mathbb{R}^m\), \[ \lim_{t=0} \frac{f \circ g(p + tV) - f \circ g (p)}{t} = df_{g(p)} \circ dg_p. \] We make use of \[ f(y) = f(x) + df_x (y - x) + R_x(y) \] with \(R_x = o(\|y-x\|)\), and \[ \lim_{t\to 0} \frac{g(p + tV) - g(p)}{t} = dg_p(V). \]
Let \(x = g(p + tV)\) and \(y = g(p)\). Then
\begin{align*} \frac{f \circ g(p + tV) - f \circ g (p)}{t} &= \frac{df_{g(p)} (g(p+tV) - g(p)) + R_{g(p)}(g(p + tV))}{t} \\ &= df_{g(p)} \left(\frac{g(p+tV) - g(p)}{t}\right) + \frac{R_{g(p)}(g(p + tV))}{t}. \end{align*}Since \(df_{g(p)}\) is a linear map \(\mathbb{R}^m \to \mathbb{R}^k\) it is in particular, continuous. Therefore \[ \lim_{t\to 0} df_{g(p)} \left(\frac{g(p+tV) - g(p)}{t}\right) = df_{g(p)} (dg_p(V)). \]
Since \(R_{g(p)} (g(p+tV)) = o(\|p + tV - p\|) = o(\|tV\|) = o(|t|\|V\|)\), we have that \[ \lim_{t\to 0} \frac{R_{g(p)} (g(p+tV))}{t} = \|V\| \lim_{t\to 0} \frac{R_{g(p)} (g(p+tV))}{t\|V\|} = 0. \]
Thus \[ \lim_{t \to 0} \frac{f \circ g(p + tV) - f \circ g (p)}{t} = df_{g(p)} \circ dg_p (V). \]
Let \(f : U \subseteq \mathbb{R}^m \to \mathbb{R}^k\) and \(g : V \subseteq \mathbb{R}^n \to \mathbb{R}^m\) be \(C^1\) functions with \(U, V\) open. Then \(f \circ g\) is \(C^1\) on the open set \(W = g^{\ast} (U)\).
By the theorem, \(d(f \circ g)_p = df_{g(p)} \circ dg_p\). Since \(f, g\) are \(C^1\), the map \(p \mapsto df_{g(p)} \circ dg_p = d(f\circ g)_p\) is continuous, hence \(f \circ g\) is \(C^1\).
Let \(g(u, v) = (u, e^{v-u}, u + v)\) and let \(f(x, y, z) = xyz\).
We have
\begin{equation*} dg = \begin{pmatrix} 1 & 0 \\ - e^{v-u} & e^{v-u} \\ 1 & 1 \end{pmatrix}, \end{equation*}and
\begin{equation*} df = \begin{pmatrix} yz & xz & xy \end{pmatrix}. \end{equation*}Thus
\begin{equation*} df_{g(u, v)} = \begin{pmatrix} (u+z)e^{v-u} & u(u + v) & u e^{v-u} \end{pmatrix}. \end{equation*}Therefore
\begin{align*} df_{g(u, v)} \circ dg_{(u,v)} &= \begin{pmatrix} (u+z)e^{v-u} & u(u + v) & u e^{v-u} \end{pmatrix} \begin{pmatrix} 1 & 0 \\ - e^{v-u} & e^{v-u} \\ 1 & 1 \end{pmatrix} \\ &= \begin{pmatrix} (u+z)e^{v-u} - u(u + v)e^{v-u} + u e^{v-u} & u(u + v)e^{v-u} + u e^{v-u} \end{pmatrix}. \end{align*}On the other hand, \[ (f \circ g)(u, v) = u(u+v)e^{v-u}. \] Thus
\begin{equation*} d(f \circ g) = \begin{pmatrix} (u+z)e^{v-u} + u e^{v-u} - u(u + v)e^{v-u} & u e^{v-u} + u(u + v)e^{v-u} \end{pmatrix}. \end{equation*}Let \(c : (a, b) \to \mathbb{R}^n\) and let \(f : \mathbb{R}^n \to \mathbb{R}\) be \(C^1\) functions. Then for \(s \in (a, b)\), \[ (f \circ c)'(s) = df_{c(s)} (c'(s)). \] Writing \(c(s) = (x^1(s), \dots, x^n(s))\) we have \[ (f \circ c)' = (x^1)' \partial_1 f \circ c + \cdots + (x^n)' \partial_n f \circ c. \]
This is just matter of tracking through the definitions. What is sometimes a little hard to wrap your head around is thinking of \(\mathbb{R}\) as a one dimensional vector space!
First, \(\mathbb{R}\) is a one-dimensional vector space with basis \(\{1\}\). Then \[ dc_s (1) = \lim_{t\to 0} \frac{c(s + t) - c(t)}{t} = c'(s). \] Note here that \(c(s + t) = c(s + t \cdot 1)\). That is, using our prior notation, \(p = s\), \(V = 1\) so that \(c(p + tV) = c(s + t)\). Similarly, \[ d(f\circ c)_s (1) = (f \circ c)'(s). \]
Then by the chain rule, \[ (f \circ c)'(s) = d (f\circ c)_s (1) = df_{c(s)} \circ dc_s (1)= df_{c(s)} (c'(s)). \]
Let \(c(t) = (\cos t, \sin t)\) and let \(f(x, y) = xy\). Then \[ c'(t) = (-\sin t, \cos t), \] and \[ df_{(x,y)} = (y, x). \] Thus \[ df_{c(t)} = (\sin t, \cos t) \] giving \[ df_{c(t)} (c'(t)) = (\sin t, \cos t) \cdot (-\sin t, \cos t) = -\sin^2 t + \cos^2 t. \]
On the other hand, \[ f \circ c (t) = \cos t \sin t \] hence \[ (f \circ c)'(t) = -\sin^2 t + \cos^2 t. \]
Let \(f : U \subseteq \mathbb{R}^n \to \mathbb{R}^k\) be a differentiable at \(p \in U\) with \(U\) open, and let \(V \in \mathbb{R}^n\). Then for any differentiable curve \(c : (-\epsilon, \epsilon) \to \mathbb{R}^n\) with \(c(0) = p\) and \(c'(0) = V\) we have \[ df_p(V) = (f \circ c)'(0). \]
For \(k = 1\), this is just the previous lemma: \[ df_p(V) = df_{c(0)} (c'(0)) = (f \circ c)'(0). \]
For \(k > 1\), writing \(f = (f^1, \dots, f^k)\) we may apply the above to the components:
\begin{align*} df_p (V) &= \big(df^1_p (V), \dots, df^k_p(V)\big) \\ &= \big((f^1 \circ c)'(0), \dots, (f^k \circ c)'(0)\big) \\ &= (f \circ c)'(0). \end{align*}Note in particular that \(c(t) = p + tV\) satisfies the requirements and this is just our original definition. The result allows us to compute the
We have seen continuous bump functions, transition functions and partitions of unity. It's possible to construct smooth versions of such functions.
The basic building block is the function,
\begin{equation*} \rho(x) = \begin{cases} e^{-1/x}, & x > 0 \\ 0, & x \leq 0. \end{cases} \end{equation*}The function \(\rho\) defined above is smooth.
For \(x \neq 0\), this is clear since \(-1/x\), \(x \neq 0\) and \(\exp y\) are smooth and \(\rho\) is a composition of these two functions via \(y = -1/x\).
For \(x = 0\), there are a couple of ways to proceed. One way is to take limits of the difference quotient \(\tfrac{\rho(x) - \rho(0)}{x - 0}\), then repeat for higher derivatives using an induction argument. Alternatively, one can show that \(\lim_{x\to 0} f^{(n)} (x) = 0\) and apply the mean value theorem to conclude that \(f^{(n)}\) exists and is equal \(0\).
Complete the proof.
From such a function, we can construct some interesting functions.
Let \[ c(x) = \frac{\rho(x)}{\rho(x) + \rho(1-x)}. \] Then \(c\) is smooth since the denominator \(\rho(x) + \rho(1-x) \neq 0\) for every \(x\). Moreover, \(c(x) = 0\) for \(x \leq 0\) and \(c(x) = 1\) for \(x \geq 1\).
Such a function is referred to as a transition function or a cutoff function . The former since the function continuously transitions from \(0\) to \(1\) as \(x\) ranges from negative values to values greater than \(1\). The latter since the function "cuts off" negative values of \(x\).
On \(\mathbb{R}^n\), the function
\begin{equation*} \rho(x) = \begin{cases} \exp\left(\tfrac{1]{1-\|x\|^2}\right), & \|x\| < 1 \\ 0, & \|x\| \geq 0 \end{equation*}is smooth.
We have that \(\rho(x) \leq 1\), \(\rho(0) = 1\), and \(\rho(x) = 0\) for \(x\) outside of the unit ball \(\mathbb{B}_1(0)\) centred on the origin.
For \(x_0 \in \mathbb{R}^n\) and \(r > 0\), the function \[ \rho_{r,x_0} = \rho(r(x-x_0)) \] satisfies \(\rho_{r,x_0} (x) \leq 1\), \(\rho_{r,x_0} (x_0) = 1\), and \(\rho_{r,x_0}(x) = 0\) for \(x\) outside of the ball \(\mathbb{B}_r(x_0)\) with radius \(r\) and centre \(x_0\).