1/18/2015

Notes of Introductory

The essay is about the nature and limits of the power which can be legitimately exercised by society over the individual.

Tyranny of the majority. Protection against the tyranny of the magistrate is not enough; there needs protection also against the tyranny of the prevailing opinion and feeling. There should be a limit to the legit interference of collective opinion with individual independence.

The question is how to find the limit.

The principle of the essay: the sole end for which mankind are warranted, individually or collectively, in interfering with the liberty of action of any of their number, is self-protection. That the only purpose for which power can be rightfully exercised over any member of a civilized community, against his will, is to prevent harm to others. Over himself, over his own body and mind, the individual is sovereign.

Notice that the above principle is meant to apply only to human beings in the maturity of their abilities. As for children, they must be protected against their own actions as well as against external injury. Liberty, as a principle, has no application to any state of things anterior to the time when mankind have become capable of being improved by free and equal discussion.

Three basic liberties:
(1) liberty of conscience, thought and feeling;expressing and publishing opinions.
(2) liberty of tastes and pursuits
(3) liberty of combination among individuals (persons combined are full age, and not forced or deceived)

The only freedom which deserves the name, is that of pursuing our own good in our own way, so long as we do not attempt to deprive others of theirs, or impede their efforts to obtain it.

ticking

In this paper, the author Lafollette argued that parenting should be licensed.

Lafollette first discussed two criteria for regulation:
(1) the behavior can be potentially harmful to others
Example: driving recklessly leads to injury and fatality, so driving is licensed
(2) competence is required
Example: lawyers, pshychiatrists, drivers need skills

In terms of parenting, bad parents can do great harm to children, so the first criterion is satisfied. As for competence, parenting does need skills and practice.

Lafollette believed that the only ways to deny the necessity of parenting regulation are to
(1) deny licensing any potentially harmful activity;
(2) deny the two criteria of regulation
(3) deny that parenting satisfies the two criteria, assuming the validity of the criteria
(4) show that even though parenting satisfies the two criteria, there exist special reasons that licensing parents isn't good
(5) show that there is no reliable and just procedure for implementing regulation

(1)(2)(3) are trivial, the author focused on (4)(5)

Theoretical objections to licensing:
(1) people have free right to have children, which shouldn't be licensed.

Lafollette's opinion: even if people have right to have children, the right might be limited to protect innocent people, i.e. children. Parents' way of rearing children might be unacceptable even if they consider it fit. He believes that a person has a right to rear children if he meets certain minimal standards of child rearing. Denying a parenting license to someone who is not competent does not violate that person's rights. Rights are conditional.

The purpose of licesing is to prevent serious harm to children, and prior restraint required by licensing would not be very onerous for many people.

(2) a parental licensing program would deny licenses to applicants judged to be incompetent even though they had never maltreated any children (prior restraint)

The author argues that individuals denied license are given the opportunity to reapply easily and repeatedly for a license. [take counseling or therapy to improve their chances of passing the next test]

The author reasoned that: even though one needs to be worry about prior restraint, if the potential for harm is great and the restraint is minor relative to the harm we are trying to prevent, then such restraint is justified. 

My thoughts:

Practical objections to licensing:
(1) what are adequate criteria of "a good parent"? Knowledge problem

Author: the proposal is only to exclude the very bad parents.

My thoughts: in terms of excluding bad parents, probably only the extreme cases are to be distinguished. What about other cases?

(2) there is no reliable way to predict who will maltreat their children

Author: Other licensing programs are not predict 100% accurately, so we cannot demand more in terms of parental regulation. We can use existing tests that claim to isolate relevant predictive characteristics that signifies bad parents to find potentially bad parents. He also believes that tests will be improved.

My thoughts: just because the other similar regulation is not perfect doesn't mean you can adopt this system into parenting.

(3) administrators might unintentionally misuse that test, which would harm innocent individuals.

Author: the fact that mistakes are made shouldn't lead us to abandon attempts to determine competence.

My thoughts: the same as before. In addition, to get pass the test, individuals might use other ways to cheat.

(4) any testing procedure will be intentionally abused.(administrators' self interest)

Author: there is no reason to believe that the licensing of parents is more likely to be abused than drivers' license tests or other regulatory procedures. In addition, individuals can appeal.

My thoughts: red-tape? bureaucrats' self-interest?

(5) It is hard to implement this program

Author: if it is important enough to protect children from being maltreated by parents, then surely a reasonable enforcement procedure can be secured.

My thoughts: costs?

1/14/2015

1/14/2015

Value theory: What is the good life? What's worth pursuing for its own sake?

Normative ethics: What're our fundamental moral duties? What counts as virtues, vices and why?

Metaethics: What's the status of moral claims and advice?

Certainly, many laws require what morality requires, and forbid what morality forbids. But the fit is hardly perfect, and that shows that morality is something different from the law. Good manners, similarly, are not the same thing as morally good conduct. The same is true when it comes to self-interest or tradition. In terms of tradition, people sometimes speak of conventional morality, which is the set of traditional principles that are widely shared within a culture or society. However, some social standards and mutual endorsement can be morally mistaken. As a result, the moral we refer to in this context are standards that are independent of conventional morality and ca be used to critically evaluate its merits. This implies the existence of some independent, critical morality that
(1) doesn't have its origin in social agreements
(2) is untainted by mistaken beliefs, irrationality, or popular prejudices
(3) can serve as the true standard for determining when conventional morality has got it right and when it has fallen into error.

How to distinguish between valid and invalid arguments.
(1) identify all of an argument's premises
(2) imagine that all of them are true
(3) ask yourself: suppose all premises were true, could the conclusion be false? If yes, the argument is invalid. If no, the statement is valid.

But a valid argument is not enough, because in testing the logical structure we assumed the truth of premises, which might not be the case. What we need is a sound argument, and that means the logic is valid and premises are true.

Example:
premise: it is morally acceptable for nonhuman animals to kill and eat other animals
conclusion: it is morally acceptable for human beings to kill and eat other nonhuman animals.

the premise doesn't logically support the conclusion. (What is morally acceptable for animals is not necessarily morally acceptable for us)

This shows that this particular way of defending is no good.

A valid argument:
premises:
(1)if it is morally acceptable for nonhuman animals to kill and eat another, then it is morally acceptable for humans to kill and eat other nonhuman animals.
(2)it is morally acceptable for nonhuman animals to kill and eat other animals
conclusion:
it is morally acceptable for human beings to kill and eat other nonhuman animals

The argument is valid but the first premise is wrong. Animals are not moral agents--they cannot control their behavior through moral reasoning, and thus it is not convincing to look to animals for moral guidance.


12/08/2014

Traditional treasury bills and TIPS

TIPS----Treasury Inflation-Protected Securities.

As its name implies, TIPS' major property is that the principal is adjusted to nation's price level, usually measure as CPI. For example, suppose the coupon rates of a ten-year TIPS and a ten-year traditional treasury bill are both 3 percent. If investors buy them at par and hold to maturity, and average inflation is 2.5 for the next 10 years, then the real return of the traditional treasury bill is only 0.5 percent while that of the TIPS is still 3.5 percent. (WHY? Because the principal coupon rate is automatically adjusted to be 2.5+3=5.5 percent) So TIPS keeps the purchasing power.

Ideally, the return yield between TIPS and traditional treasury bill is the average market participants' inflation expectation, which can give us valuable information in inflation forecasting. In real life, however, some other factors also play roles in this yield spread.

(1) Inflation risk:
TIPS is inflation-risk free, while traditional treasury bill is subject to inflation effect, so investors bear more risk buying the latter bond. Based on economic theory, investors should be compensated by some amount to buy the bond, and we call this amount of pecuniary compensation as Inflation Risk Premium.

(2) Liquidity risk:
Based on finance theory, value of a certain type of asset is related to its liquidity. If the asset takes significant time and effort to sell at fair price, then investors might have to lower price to sell it when time constraint is tight. Traditional treasury bills are supposed to be the most liquid assets in financial market, and thus investors of these bonds don't have such worries. TIPS are relatively more illiquid and thus liquidity premium might exist to attract investors to buy TIPS. (TIPS have higher liquidity risk)

To summarize, the yield spread can be expressed as:

y^(n) - y^(r) = inflation expectation + inflation risk premium - liquidity risk premium,

where  y^(n) denotes nominal return of traditional treasury bills, and y^(r) denotes real return of TIPS.

From this equation it is clear that the spread conveys accurate info about inflation expectation only when inflation risk premium is of the same size as liquidity risk premium. However, if both premium are significant smaller than inflation expectation, then the spread still has a good approximation of inflation expectation. Furthermore, if premiums are roughly constant over time, then we can know change of inflation expectation by tracking change of spread rate.






12/06/2014

Paper review of Stock and Watson's Forecasting Inflation

The paper I review is John Stock and Mark Watson's \textit{Forecasting Inflation}, which was published in 1999. The main focus of the paper is on Phillips curve's forecast ability, and authors aimed to answer the four following questions. Has the curve been stable over time? Has the unemployment rate-based Phillips curve done a good job predicting U.S. inflation at 12-month horizon? Is it possible that the curve can be improved by incorporating some other real economic activity variables? If not, does there exist any other model that we can resort to? 
 
First introduced by British economist William Phillips in 1958, the Phillips curve shows that rising price level helps decreasing unemployment rate. However, this inverse relationship doesn't hold in the long run because market participants tend to adjust their inflation expectation and thus as time goes on price level increases while unemployment rate returns back to its natural level. (Friedman, 1968) So in terms of forecasting inflation, researchers typically limit prediction horizons to 12-month so that the inverse relationship exists.
 
In their paper, Stock and Watson (abbreviated as S\&W in the rest of the paper) broadly interpreted Phillips curve as short-run relationship between future inflation and current real market activity. In other words, they aggregate variables in addition to unemployment rate in regressions and examined if they help improve forecast performance of the traditional Phillips curve. To discover existence of better alternative prediction tools, S\&W also considered models either backed by economic theories or more sophisticated in model construction. S\&W used U.S. monthly data from 1959 January to 1997 September on inflation as measured by CPI and Personal Consumption Expenditure (PCE) deflator. The data source is DRI-McGraw Hill Basic Economics Database.
 
The base model S\&W used is as follows: $$\pi^{h}_{t+h}-\pi_{t}=\phi+\beta(L)u_{t}+\gamma(L)\Delta\pi_{t}+e_{t+h}$$ where 
 
\begin{enumerate}
\item $\pi^{h}_{t+h}=\frac{1200}{h}\ln (\frac{P_{t+h}}{P_{t}})$: h-period ahead inflation in the price level $P_{t+h}$, reported at an annual rate
\item $\pi_{t}=1200 \ln(\frac{P_{t}}{P_{t-1}})$: monthly inflation at annual rate
\item $u_{t}$: unemployment rate
\item $\Delta\pi_{t}$: first order difference of inflation
\item $\beta(L)$, $\gamma(L)$: polynomials in the lag operator L 
\end{enumerate}
 
There are two assumptions in the model: inflation data has a unit root and the natural rate of unemployment (NAIRU) is constant.\footnote{Indeed, the natural rate $\bar{u}$ is captured by the constant term $\phi$}. These two assumptions are common in related literature but if future researches find strong evidence against them, then the validity of S\&W's tests are questionable.
 
S\&W first tested whether coefficients in the model have been stable over the sample period. They used Quandt likelihood ratio (QLR) test to discover unknown breakdate. The null hypothesis of QLR(all) is that all coefficients are jointly stable; QLR($\phi,\beta$) tests whether $\phi$ and $\beta(L)$ are jointly stable, assuming constancy of $\gamma(L)$, and QLR($\gamma$) is just the opposite. Table.1 shows their testing result. QLR($\gamma$)s have small p-values, implying that this variable is unstable, and the structural break is 1983. However, QLR test may not be the best testing method, and there could exist more than one structural break. Even though they found model was unstable, S\&W ignored it due to quantitative insignificance of $\gamma(L)$ and regarded it as the benchmark model. 
 
S\&W broadly interpreted Phillips curve as short run relationship between future inflation and current real market activities. The generalized model is as follows: $$\pi^{h}_{t+h}-\pi_{t}=\phi+\beta(L)x_{t}+\gamma(L)\Delta\pi_{t}+e_{t+h},$$ where $x_{t}$ denotes real market variables other than unemployment rate and is either processed or assumed to be stationary. Their detrending method is one-sided Hodrick-Prescott (HP) filter. 
 
 They first used recursive method to estimate in-sample model coefficients and then conducted pseudo out-of-sample forecast. To compare these models' forecast performance, authors used relative mean square error (Rel.MSE) with respect to benchmark Phillips curve and forecast combination regression method. Namely, they regressed the following model: $$\pi^{h}_{t+h}-\pi_{t}=\lambda f^{X}_{t}+(1-\lambda)f^{U}_{t}+\epsilon_{t+h},$$ where 
\begin{enumerate}
\item $f^{X}_{t}$ denotes forecasts based on candidate variable X, made at time t
\item $f^{U}_{t}$ denotes forecasts based on unemployment rate, made at time t
\item $\lambda$ shows how much forecast based on x helps improve forecast ability of benchmark unemployment forecast. 
\end{enumerate}
 
If Rel.MSE is statistically less than 1, then it means its corresponding forecast is better than the benchmark. If $\lambda$ is statistically positive, then the corresponding variable adds forecast accuracy to the benchmark based on unemployment rate.
 
Table.2 and Table.3 tabulated their results. We can observe that in general PCE inflation forecasts are most accurate than CPI inflation ones. Forecast results in the second-half period (1984-1996) are more accurate than the first half. What's more, over years the variables \textit{capacity utilizations} (i.e.IPXMCA) and \textit{manufacturing and trades sales} (i.e.MSMTQ) both have small Rel.MSEs and positive $\lambda$s. So these two variables  help improve forecast that is based on unemployment rate. 
 
To discover potential existence of better models, S\&W also considered models backed by term structure expectations, nominal money supply theory, and multivariate models incorporated with different kinds of variables.\footnote{S\&W used the same methods of coefficient estimation and forecast performance comparison. Since the result tables are too large, I choose not to include them in the paper.} They found that: (1) bivariate models based on term structure of interest rates or nominal money supply produced worse forecast results than Phillips curve; (2) bivariate models incorporating exchange rate or different price indexes didn't outperform Phillips curve consistently; (3) multivariate regressions showed that incorporation of other real activity variables didn't significantly improve the forecasting results. 
 
Based on their findings, S\&W concluded that Phillips curve produced the most reliable short-run inflation forecast among all models they considered and certain aggregate activities variables, namely capacity utilization and manufacturing and trade sales help improve the curve's forecast ability. Their conclusions echoed with some previous researches. For example, Alan Blinder (2001), the former vice chairman of the Board of Governors of the Federal Reserve System, claimed that the curve has done amazingly well over years and it deserves a prominent place in models used for policy-making purposes.  
 
However, there are certain drawbacks in this paper. First of all, since the authors used a large number of forecasts, overfitting bias is inevitable. In addition, they assumed that inflation is I(1), which is controversial. If future researches evidently prove that inflation is integrated of other orders, then this paper's findings might not be valid. Furthermore, S\&W only considered linear models for the sake of simplicity, but as they admitted, the relationship between inflation and real economy variables can be much more complicated. Indeed, Ang, Bekaert and Wei (2006) argued that this paper failed to consider non-linear models like no-arbitrage term structure models or various survey data and thus their findings are incomplete and unconvincing. The harshest critique, however, comes from Atkeson and Ohanian (2001). They used the following random walk model as the benchmark  $$\pi^{12}_{t+12}=\pi^{12}_{t}+v^{12}_{t+12},$$ where $\pi^{12}_{t+12}$ is the forecast of the 12-month rate of inflation and $v^{12}_{t+12}$ is white noise, and used the same forecasting comparison methods as S\&W did. Surprisingly, they found that from 1985 to 2000, none of S\&W's models performed better than this naive random-walk model. This suggests that the best indicator of future inflation is current price level, which, if true, makes researchers' models based on sophisticated theories pretty futile. However, as S\&W commented in their later paper (2008), the robustness of Atkeson-Ohanian model depends delicately on sample periods and forecast horizon. 
 
Forecasting inflation is hard and Phillips curve is merely one candidate model. Recently some researchers (Spen and Corning (2001), Calstrom and Fuerst (2004)) start to use Treasury Inflation-Protected Securities (TIPS) for price level prediction. As its name implies, TIPS are the inflation-indexed bonds whose principal is adjusted to the CPI. If CPI changes, the principal adjusts correspondingly to maintain the same purchasing power. By studying behavior of TIPS, we can gain information on investors' expectation of inflation and do forecast based on that.
 
References
\item Friedman, Milton, "The Role of Monetary Policy", American Economic Review, March 1968, pp 1-17
 
\item Blinder, Alan, "Is There A Core of Practical Macroeconomics That We Should Believe?", American Economic Review, May 1997, pp 240-243.
 
\item Ang, Andrew, Green Bekaert and Min Wei, "Do Macro Variables, Asset Markets, or Surveys Forecast Inflation Better?" Journal of Monetary Economics, May 2007, pp. 1163-1212
 
\item Atkeson, Andrew and Lee Ohanian, "Are Phillips Curve Useful for Forecasting Inflation?", Quarterly Review, The Federal Reserve Bank of Minneapolis, 2001
 
\item Stock, John and Mark Watson, "Phillips Curves Inflation Forecasts", NBER Working Paper 14322, September 2008
 
\item Spen, Pu and Jonathan Corning, "Can TIPS Help Identify Long-term Inflation Expectations?", Federal Reserve Bank of Kansas, Economic Review, Fourth Quarter, 2001
 
\item Carlstrom, Charles and Timothy Fuerst, "Expected Inflation and TIPS", Federal Reserve Bank of Cleveland, November 2004

Tables:



11/09/2014

Fibonacci Sequence (LaTex file)

In mathematics, Fibonacci sequence is a sequence of positive integers 

$$a_{1}, a_{2}, a_{3},\cdots a_{n} $$

which satisfies the following conditions

\begin{enumerate}
\item $a_{1}=1$,
\item $a_{1}=1$,
\item $a_{n}=a_{n-1}+a_{n-2}, \forall n \geqslant 2$
\end{enumerate}

This is a recurrence relation and we need to find out a general formula to denote every single member of the sequence. 

There are many ways to solve for this question. For example, a combinatorial expert may define a generating function $F(x)=\sum^{\infty}_{n=1} a_{n}x^{n}$ where $a_{i}$ is a Fibonacci number.\footnote{For a detailed proof of this method, refer to Richard Brualdi's \textit{Introductory Combinatorics}, 5th edition.}  By solve for $F(x)$ explicitly, we can write out the general formula of the sequence: $${a_{n}=(\frac{1}{\sqrt{5}})(\frac{1-\sqrt{5}}{2})^{n}-\frac{1}{\sqrt{5}} (\frac{1+\sqrt{5}}{2})^{n}}$$.

Another widely recognized proof uses matrix. By noticing the recurrence relation is a linear difference equation, we can set up a two-dimensional system
\[
                \left(   \begin{array}{r} a_{n+2} \\ a_{n+1}  \end{array}\right)= \left(   \begin{array}{rr} 1 & 1 \\ 1 & 0  \end{array}\right)\times \left(   \begin{array}{r} a_{n+1} \\ a_{n}  \end{array}\right).
\]
Using eigenvalues and respective eigenvectors of this system, we can also get the same general formula of the sequence. \footnote{For a more detailed proof, refer to Henry Gould (1981)}


In this post, I will use another method to prove the sequence's general formula. The methodis conceptually straightforward and requires no further knowledge than basic geometric sequence properties and undetermined coefficient calculation. In specific, I will construct a couple of geometric sequences from the relation $a_{n}=a_{n-1}+a_{n-2}$, calculate their general formulas, and combine them to get the final result. 

First of all, we know that for a geometric sequence {$b_{n}$}, if $b_{1}=k$ and its recurrence relation is $b_{n}=mb_{n-1}$, ($k \in \mathbb{R}, m \in \mathbb{R}$), then we know $b_{n}=k m^{n-1}$.
Given the recurrence relation $a_{n}=a_{n-1}+a_{n-2}, n \geqslant 3$, it is possible to rewrite this equation and construct a geometric sequence, and by writing out its general formula, we are one step closer to calculating the general form of Fibonacci numbers. 
Rewrite   $a_{n}=a_{n-1}+a_{n-2}$ to be $a_{n}+y(a_{n-1})=x(a_{n-1}+ya_{n-2})$, and this implies that if we can find real numbers $x$ and $y$ to make the two equations identical, then the sequence ${a_{n}+ya_{n-1}}$ is a geometric sequence with common ratio $x$. 

After some calculations, we can get $a_{n}+y(a_{n-1})=x(a_{n-1}+ya_{n-2})$ is $a_{n}=(x-y)a_{n-1}+xya_{n}$, and thus it is obvious that \begin{subequations}
\begin{align}
        x-y &= 1,\\
        xy &= 1,
\end{align}
\end{subequations}
Thus, $x=\frac{-1 \pm \sqrt{5}}{2}$ and $y=\frac{1 \pm \sqrt{5}}{2}$. We choose $x$ to be $\frac{1+\sqrt{5}}{2}$, and $y$ is $\frac{-1+\sqrt{5}}{2}$ correspondingly \footnote{Choosing the other bundle yields the same final result.}.
Now that $a_{n}+\frac{-1+\sqrt{5}}{2}(a_{n-1})=\frac{1+\sqrt{5}}{2}(a_{n-1}+\frac{-1+\sqrt{5}}{2}(a_{n-2}))$, we can see that $a_{n+1}+\frac{-1+\sqrt{5}}{2}(a_{n})$ is a geometric sequence whose initial term is $1+\frac{-1+\sqrt{5}}{2}=\frac{1+\sqrt{5}}{2}$ and common ratio is $\frac{1+\sqrt{5}}{2}$. So  $$a_{n+1}+\frac{-1+\sqrt{5}}{2}(a_{n})=(\frac{1+\sqrt{5}}{2})^n.$$
Seeing that the strategy of constructing geometric sequence works in the first stage, we might be attempted to do it again to the above recurrence relation. However, it is important to note that the right hand side term is no longer a constant; instead, it is also related to $n$. Thus, we need to add rewrite the relation slightly differently. 
Let $a_{n+1}+\lambda (\frac{1+\sqrt{5}}{2})^{n+1}=\frac{1-\sqrt{5}}{2}(a_{n}+\lambda (\frac{1+\sqrt{5}}{2})^{n})$, and compare it to the original equation $a_{n+1}+\frac{-1+\sqrt{5}}{2}(a_{n})=(\frac{1+\sqrt{5}}{2})^n$, we know to equalize the two equations, it must be the case that $$\lambda (\frac{1-\sqrt{5}}{2}) (\frac{1+\sqrt{5}}{2})^{n}-\lambda (\frac{1+\sqrt{5}}{2})^{n+1}=(\frac{1+\sqrt{5}}{2})^{n}$$
After some computation work, we can get that $\lambda=-\frac{1}{\sqrt{5}}$. Plug it back in the recurrence relation, we get $$a_{n+1}-\frac{1}{\sqrt{5}} (\frac{1+\sqrt{5}}{2})^{n+1}=\frac{1-\sqrt{5}}{2}(a_{n}-\frac{1}{\sqrt{5}} (\frac{1+\sqrt{5}}{2})^{n})$$

Although it looks a little bit ugly, actually the sequence ${a_{n}-\frac{1}{\sqrt{5}} (\frac{1+\sqrt{5}}{2})^{n}}$ is a geometric sequence with initial term $\frac{5-\sqrt{5}}{10}$ and common ratio  $\frac{1-\sqrt{5}}{2}$. So $${a_{n}-\frac{1}{\sqrt{5}} (\frac{1+\sqrt{5}}{2})^{n}}=(\frac{5-\sqrt{5}}{10})(\frac{1-\sqrt{5}}{2})^{n-1}$$. 

Note that $\frac{5-\sqrt{5}}{10}=(\frac{1}{\sqrt{5}})(\frac{1-\sqrt{5}}{2}).$ If we rewrite the equation, then $${a_{n}=(\frac{1}{\sqrt{5}})(\frac{1-\sqrt{5}}{2})^{n}-\frac{1}{\sqrt{5}} (\frac{1+\sqrt{5}}{2})^{n}}$$
and this is the exact general formula of Fibonacci sequence. 

My method stems from realization that potential geometric sequences might be derived from the given recurrence relation, and since we are more familiar with geometric sequence's properties, we can use it to tackle more complicated problems. The method is conceptually naive and requires the original relations to be simple. Otherwise, we may resort to other techniques like generating functions.

REFERENCE
1. Brualdi, Richard, \textit{Introductory Combinatorics, 5th edition}
2. Gould, Henry \textit{A History of the Fibonacci Q-Matrix and a Higher-dimensional Problem}, Fibonacci Quart., 19(1981), 250-257

10/15/2014

Fermat's Last Theorem

The Fermat's Last Theorem states that for $a, b, c \in \mathbb{Z^+}$ and $n \in \mathbb{N}, n > 2$, the equation $ a^n + b^n = c^n$ doesn't hold. 

Despite its apparent simplicity in terms of hypothesis and conclusion, this theorem is actually very involved and remained a mere conjecture until Andrew Wiles formally proved it in 1995, more than 300 years after Fermat first wrote it in the margin of Diophantus's \textit{Arithmetica}. 

The page of \textit{Arithmetica} that motivated the question was about Pythagorean theorem, which states that in any right triangle, the square of hypotenuse is equal to the sum of square of two other sides. In other words, $\exists a,b,c \in \mathbb{Z^+}$ such that $a^2+b^2=c^2$.   

Having tested various solutions to the Pythagorean theorem, Fermat went a step further to check whether the theorem holds for $n=3$. Surprisingly, Fermat argued that such system is insoluble and more generally, there is no positive integer solutions to any n that is larger than $2$. He commented that he came up with a remarkable proof to solve for all cases but the margin is too narrow for him to write down.\cite{citationkey1} As a result, we are only able to read Fermat's proof for a special case where $n=4$, while the general theorem remained unsolved for the next three centuries. 

Mathematicians at first tried to find existence of counter-examples but failed, which incentivized more people to believe that Fermat's conjecture is true. Later, mathematicians managed to prove the conjecture when $n=3, 5$ and $7$, but no one was able to come up with a general proof for all equations. 

The longer one conjecture remained mysterious, the more attractive it seemed to people who inspired to solve it. One story says that in nineteenth century a German physician and mathematician Paul Wolfskehl bequeathed 100,000 marks (quite a big amount of money at that time) to the first one who could prove it. Such lucrative prize motivated more people, especially amateurs to hand in their proofs, but unfortunately none of their attempt was successful. Due to advent of modern computers after WWII, mathematicians were able to conclude that Fermat's conjecture holds for finitely large $n$, but still, nobody was able to give a complete proof. The conjecture also became notorious for its difficulty and before Wiles' proof, it was described as "most difficult mathematical problems" in the \textit{Guinness Book of World Records}. \cite{2}

In the twentieth century, Shimura and Taniyama, two Japanese mathematicians specialized in elliptic curves, conjectured that each elliptic curve can be matched to a modular. It seemed at first Shimura-Taniyama conjecture was unrelated to Fermat's conjecture, but surprisingly, Kennith Ribet and Gerhard Frey reasoned that Shimura-Taniyama conjecture doesn't hold unless there exists no positive integer solutions to $a^n+b^n=c^n$ for infinitely large $n$. In other words, Shimura-Taniyama conjecture begs the question that Fermat's last theorem is true. Hence, the Fermat's conjecture returned to mainstream of math proof not just for intellectual interest, but also for paving way for new branches of math research.   
\


Andrew Wiles, who has been obsessed with Fermat's Last Theorem since childhood, specialized in elliptic curves both as a graduate student in Cambridge and as a professor in Princeton. On hearing Ribet and Frey's research progress, he decided to prove Fermat's conjecture by attacking Shimura-Taniyama conjecture. He spent seven years working alone, learning and extending some new methods at that time,\cite{3} and presented his proof in 1993. While Wiles enjoyed public accolade and fame at first, a team of referees found in his paper a big flaw in a bound of the order of a specific group. \cite{4} Wiles spent one more year trying to repair his proof. Fortunately, before he wanted to give up an intuition struck him and help him correct the previous approach. Wiles published his manuscripts in 1995 on \textit{Annals of Mathematics}, and he proved Fermat's Last Theorem by techniques seemingly unrelated to number theory. After over 300 years of effort, mathematicians finally were able to prove Fermat's Last Theorem, and the process of proof is roundabout and interdisciplinary.