Friday, June 17, 2016

I haven't been doing nothing

It's been a year since I blogged and one might understandably have thought the whole enterprise had been given up. To the contrary, I decided that it is far better to spend weeks and months mulling over a subject and then prepare careful notes rather than thinking about things piece meal and then writing up notes. The latter way of doing things I found gave a false finality to the whole affair and produced a questionable notion of progress. The fruits of my labor will be appearing in the coming days, weeks and months on a new website I created at physicalmusings.neocities.org.
  I have prepared around 80 pages worth of latex notes on Differential Geometry and about 70 pages of  latex notes on group and representation theory. I am yet to start writing up the notes on Quantum Field Theory but I shudder when I think of the task.  On the other hand, I am quite happy with how far I have got with the tasks at hand. Understanding QFT is not very easy and the task is compounded by the fact that every author has their own way of not just arranging the material but whatever they consider primary and foundational. The result is that having a number of books on QFT unlike any other subject only served to confuse and perplex.  How to learn Quantum Field Theory is a topic that needs to be carefully considered and needs its own blog post.

Sunday, January 25, 2015

A small note on Entropy in Information Theory and Entropy in Physics

It is rather an amazing fact that if one asks for a function that will capture how much "choice" there is in selection of an event (to use Claude Shannon's term) we arrive at an equation that is exactly the one called entropy in statistical mechanics i.e \( S= - k \sum_i p_i \ln p_i\) where k is some constant. So the obvious question then is that, is this a coincidence or something deep? We have spent a bit of time carefully go through the thermal physics and ultimately arriving at entropy and drawing conclusions from the principle of increasing entropy and all throughout was the ever present notion of a quasi-static process because we always wanted to stay in equilibrium in order that we may use state functions.  This clearly has no direct relation with the entropy in thermal physics but perhaps thermal physics is the wrong place to look for the relation.  Obviously, we are led to invoke some notions from statistical mechanics.

Suppose we have a system A with entropy \( S_1\) and add then we have another system B with entropy \( S_2\). If we consider the composite system, we have that the combined entropy will be  \( S_1 + S_2\). Now concurrently, we can consider the number of states (however we define that)available to each system. Then we will have \( \Omega_1 \) states for system A and \( \Omega_2 \) states for system B. How then do we get the total number of states for the combined system? The answer is \( \Omega_1 \Omega_2 \). If we are to look for a relation between the two concepts it will have to obey

\begin{equation}
 S(A) + S(B)=  f( \Omega_A \Omega_B ) \hspace{10mm} \text{eq.1}
\end{equation}
where f is some undetermined function.

Let us for the moment switch gears to information theory. Say I have some random variables X and Y which can take on values \( x_i\) and \(y_i \) respectively. We now want some function ,I, that characterizes the amount of "choice" I have from a specific probability event. Clearly this function must satisfy some intuitive properties. I only state the property that is relevant for this discussion. If I have two different probabilistic events \( x_i\) and \(y_i \), then surely we want the property that the amount of choice of both these two events will be  \( I(x_i) + I(y_i)\) and this will be the amount of information \( I(x_iy_i) \) for the observed outcome \( x_i y_i\). So we have that

\begin{equation}
 I(x_i y_i)=   I(x_i) + I(y_i) \hspace{10mm} \text{eq.2}
\end{equation}
 .
It should be clear now why in both cases the desired function turns out to be natural log.
In both situations we are solving Cauchy's functional equation but can we go deeper? As far as I can tell the answer is no(so far). No one as far as I can tell has come up with a conceptual understanding that ties the two areas together. People use the word entropy in both situations, it is the same equation but this does not imply that there is the same concept at work. Notice that the formula for entropy looks like an average where the function being averaged over is the natural log. So we are calculating the average information over some probabilistic distribution but here comes the crucial point; who said it had to be the Boltzmann distribution or any distribution relevant for physics? The differential equation for the simple harmonic oscillator appears in very disparate situations but then no one argues for a connection between the situations.

There is a paper by E.T Jaynes called \( \textit{Information Theory and Statistical Mechanics} \) claiming to make the connection. I am not sure I buy it since it merely takes advantage of the fact that the equation is the same in both cases and goes on from there. In fact he goes on the derive the Boltzmann distribution doing exactly the same kind of calculation done in Statistical mechanics the difference is how he interprets the calculation. He interprets it in such a way that one never needs to consider any physical assumption! To quote  part of the abstract:

" It is concluded that statistical mechanics need not be considered as a  physical theory  dependent for its validity on the truth of additional  assumptions not contained in the laws of mechanics (such as ergodicity, metric transitivity, equal a priori probabilities etc)''

In other words if experiments falsified his predictions(which they don't), his interpretation would still stay where it is namely, this is provably what you get when you consider statistical inference. It just so happens that this statistical inference (containing no physical assumptions or necessities) corresponds to reality. So if he takes out the physics from statistical mechanics why should we conclude that he has found a connection between information theory and physics?

Clearly there is something deep happening but again no one has gone further than commenting on the fact that the same equation appears in both fields. As a result we can draw no conclusion.


Saturday, January 24, 2015

Principle of Increasing Entropy, consequences and Maxwell Relations

Note that I am not calling the principle of increasing entropy the second law of thermodynamics because it was a consequence two previous states that I called the second law of thermodynamics. This can all be thrown aside as semantics that does not contain any invariant information. Setting aside this debate, we can move on to derive the consequences.
Let us first note, that the first law of thermodynamics however powerful it is, is not directly relate-able to experiments, after all it says that \(U = TdS - PdV,\) but how are we going to construct a knob that we can use to control entropy. We therefore need other kinds of state functions that also have units of energy and can easily be transferred to experiments. We shall introduce these functions by taking legendre transforms of U.

First let us introduce one that is important to chemists but not for physicist.
Consider the function \(H= U +PV\). We now take the differential of both sides to get \(dH = dU + VdP + PdV\) but we know what dU is from the first law, so we that
\(dH = TdS -PdV + VdP +PdV = TdS + VdP \) so H(S,P).  Now, we recall that the function we are dealing with is a state function and therefore does not depend on the path in otherwords it is an exact differential.
\(F = AdX + BdY \) is exact when \( \frac{\partial A}{\partial Y})_X= \frac{\partial B}{\partial X})_Y\). We apply this to H to get \( \frac{\partial T}{\partial P})_S =  \frac{\partial V}{\partial S})_P \). This is our first example of Maxwell Relation, derived painlessly without jumping into a morass of partials. Notice the natural variables of the state function are those that are held constant on each side the relation. This is how one can tell which state function to use in order to derive the relation.

Our second state function is important to physicists and that is got by taking the following legendre transform \(F= U - TS\). We play the same game as before and take the differential of both sidesto get that  \(dF = -SdT - PdV\). This is an exact differential so \( \frac{\partial S}{\partial V})_V=  \frac{\partial P}{\partial T})_T \)

The third example is got by considering \(G = F+PV\). This gives a differential of
\(dG = -SdT + VdP\). Applying the condition for exact differential gives  \( \frac{\partial S}{\partial P})_T = - \frac{\partial V}{\partial T})_P \). The reader can do the same exercise for the internal energy U.

The purpose of these other state functions is to ease analysis of equilibrium states. If one has control over the temperature and volume of a system and considers those as the important variables then one should not insist with working with the internal energy but instead work with F or the Helmholtz free energy or if the thermodynamic coordinates are T and P the natural state function to consider is the the Gibbs free energy, G.

Let us now consider two situations which one is likely to come upon in thermodynamics and we shall use the principle of increasing entropy and the first law of thermodynamics to draw important conclusions.
Let there be a system coupled with a thermal reservoir at  temperature \( T_0\). We shall assume that the system can't expand in volume, but there might be some internal degrees of freedom of the system that are not in equilibrium that allow for heat, Q, to be transferred to the reservoir. These situation might represent a chemical reaction or may a cup of ice and water sitting on table. The cup is the system and surrounding is the reservoir at temperature \( T_0\). Let us further assume that the system and the reservoir make up our universe i.e they both are thermally isolated. We know from the principle of increasing entropy that
 \begin{equation}
    \Delta S + \Delta S_0 \geq 0  \hspace{10mm} \text{eq.1}
\end{equation}.
S, here is the entropy of the system and \( S_0\) is the entropy of the reservoir. We know that \( S_0 = - \frac{Q}{T_0}\) and plugging this into eq.1 we get
\begin{equation}
Q - T_0  \Delta S \leq 0  \hspace{10mm} \text{eq.2}
\end{equation}
We know from the first law that \( \Delta U = Q \), since there is no change work being done by the system and plugging this into eq.2 we get
\begin{equation}
\Delta(U - T_0 S ) = \Delta F \leq 0  \hspace{10mm} \text{eq.3}
\end{equation}
Therefore for a system coupled to a thermal reservoir at fixed volume minimizes the free helmholtz energy if we obey the principle of increasing entropy. So as the system achieves thermodynamic equilibrium F goes to a  minimum.

Let us consider another situation where we have a system that can do work that is calculated by \(P_0 \Delta V\) (this is the only work we will consider). This time the reservoir is a thermal reservoir at temperature  \( T_0 \) and a pressure reservoir with pressure at \( P_0 \). Again we will there will be some heat exchange and we apply the principle of increasing entropy i.e eq.1 and follow by the same arguments as before to eq.2. This time we know from the first law that \( \Delta U = Q - P_0 \Delta V\). Solving for Q and plugging into eq.2 we arrive at

\begin{equation}
\Delta (U + P_0V - T_0 S) = \Delta G \leq 0 \hspace{10mm} \text{eq.4}

\end{equation}

But this is precisely the Gibbs free energy. So the principle of increasing entropy implies for a system coupled to a thermal and pressure reservoir that Gibbs Free energy is minimized.

Eq.3 and Eq.4 apply for both reversible and irreversible processes since state functions do not care about history or path but simply the endpoints which are equilibrium states.
 







Wednesday, January 21, 2015

What is to come?

What is usually described as the second law of thermodynamics, namely that entropy of an isolated system stays constant or increases as has been demonstrated is a result of more fundamental statements that have been discussed. But now that we know entropy always increases, do we stop here? Surely, this has some consequences. In the next few discussions we shall investigate what results as a consequence of the never decreasing entropy.
After which we shall leave the world of equilibrium processes and investigate the world of non-equilibrium processes. We shall not go over the necessary basic statistical mechanics as excellent resources (in my opinion) are readily available.

Friday, January 16, 2015

Clausius Inequality and Entropy

The  Clausius Inequality arises when we consider cyclic processes. This leads to the concept of entropy. For this discussion recall that
\( \frac{Q_2}{Q_1}= \frac{T_2}{T_1} \). In particular \( \delta Q_2 = \frac{T_2}{T_1} \delta Q_1 \).  Now imagine there is an engine doing work in a cyclic process, this will be done with help of carnot engines. There ill be one ultimate reservoir from where we will get our supply of hear. The engine will be taken from one temperature \( T_i\) to \( T_{i+1} \) by an auxillary reservoir at temperature \( T_i \) to keep the reservoir repleished it gets a supply of heat \( \delta Q_i \) for a carnot engine \( C_i\) but this engine gets it heat from another auxillary reservoir in the amount \( \delta Q_i \frac{\tilde{T}}{T_i} \) and this auxillary reservoir got its heat \( \delta Q_i \frac{\tilde{T}}{T_i} \) from the big reservoir at temperature \( \tilde{T} \). We use the index because movement from one temperature \( T_i \) to \( T_{i+1} \) requires a carnot engine \(C_i \). Everything is done in a cycle so that ultimately \( \Delta U = 0\). The total heat transferred to the engine in the cycle in \( \sum \frac{\delta Q_i \tilde{T}}{T_i} = Q \).

From first law we have that \( Q=W \). But this is a violation of second law because we have converted all the incoming heat into work. This can be avoided if Q and W are negative i.e all the work we put in is turned into heat. Therefore we have that \( Q=W \leq 0 \) so \( \tilde{T} \sum_{i} \frac{\delta Q_i}{T_i} \leq 0 \).

For infinitesimal \( \delta Q_i \) we have that
\begin{equation}
 \oint \frac{dQ}{T} \leq 0 \hspace{10mm} \text{This is the Clausius Inequality.}
 \end{equation}
 But we can reverse the direction of the cycle to get
\( \int \frac{dQ}{T} \geq 0 \). Therefore for a reversible process \begin{equation}
\oint \frac{dQ}{T}= 0  \hspace{10mm} \text{Reversible process}
\end{equation}

Now for a reversible process consider going in a cycle from initial state to final state and back to the initial state. We this we have
\( \int_i^f \frac{dQ}{T} + \int_f^i \frac{dQ}{T} = 0 \). So we have that
\( \int_i^f \frac{dQ}{T} = \int_i^f \frac{dQ}{T} \)
This implies the existence of a state function where \( S_f - S_i\).

We have finally arrived at the concept of entropy. Now we can show that the change in this state function always increasing or stays the same. So consider a process we go from some initial state to final state by some irreversible process and go back to the initial state by a reversible process. We therefore have
\begin{equation}
 \int_{i(irreversible)}^f \frac{dQ}{T} + \int_{f(reversible)}^i \frac{dQ}{T} \leq 0
\end{equation}
so
\begin{equation}
 \int_{i(irreversible)}^f \frac{dQ}{T} \leq \int_{i(reversible)}^f \frac{dQ}{T} = \Delta S
\end{equation}
 The first integral was over an irreversible process while the second was over a reversible process. So we have that the first integral over the irreversible process is
\begin{equation}
\int _{i(irreversible)}^f \frac{dQ}{T} \leq \Delta S \hspace{10mm} \text{Integral over an irreversible process}
\end{equation} .

Thermally isolate system we have dQ=0 so \( \Delta S \geq 0 \). This is what people call the second law of thermodynamics.



Friday, January 9, 2015

Equivalence of the two statements of the second law.

We know prove the equivalence of the two statements of the second law of thermodynamics. Recall that they are the following:

1. No engine or process exists such that its sole purpose is to turn all the heat it extracts from a reservoir into work. It must release some energy into a colder reservoir.

2. No engine or process exists whole sole purpose is to transfer heat from a cold reservoir to a hot reservoir with no other effects.

We shall prove their equivalence but assuming the violation of one of them and showing that implies the violation of the second.

Proof
Let us assume the violation of the first statement. This means that the engine E get Q from a reservoir and produces work, W, such that Q=W. Now imagine a composite system with this engine E and a second engine C. Let the work done by engine E be put in the engine C so that C can get heat \( Q_2\) from a cold reservoir and supply  \(Q_2 +W \) to the hot reservoir. So in terms of the composite system (E and C) we got Q from the hot reservoir and used engine C to put back \( Q_2 +W = Q_2 +Q_1\). Therefore we transferred \( Q_2\) from cold reservoir to the hot reservoir with no work supplied i.e from the view of the composite system, the process consisting of the engine E and engine C, the second statement of the second law was violated.

For the second part of the proof, we assume that the second statement was violated and the goal will be to show that this implies the violation of the first statement. This time assume that engine E takes  \( Q_2\) from a cold reservoir and places it all in the hot reservoir with any work supplied and with no other effects. Now imagine another engine C getting \( Q_1\) from the hot reservoir doing work, W ,and leaking   \( Q_2\) into the cold reservoir. Thus the amount of work done by engine C is \( W = Q_1 -Q_2 \). Again let us look at these two engines as one system. From this point of view . Thus the total amount of heat got from the hot reservoir was  \( Q_1 -Q_2\) and we used all of it as work. This violates the first statement.

Hence the two statements are equivalent.

Note: The first statement does not forbid us from putting work, W, into a system and the engine turning all of this into heat, Q, which is then dumped into a reservoir. This process is what will be used to arrive at the concept of entropy.

Tuesday, January 6, 2015

Carnot Engines and the Second Law of Thermodynamics

We now come to discuss two statements of the second law of thermodynamics which we shall see gives us a upper bound to the efficiency of any imaginable engine used to do work.
We start with a few experimental facts that were known by Carnot who did the original work on this topic.

Facts
1. We can get work from an engine if it is working between two heat sources of different temperatures. Clearly, we would like our engine to leave our heat sources unchanged and we would like the engine to return to its original state after a while in order to begin another cycle. These two requirements can be satisfied if our engine uses processes that are reversible.
2.It is possible for no work to be done when heat moves from a hot body to a cold body while the system is returning to equilibrium. Therefore any return to thermal equilibrium where no work has be done must counted as loss. We therefore want the engine to work between heat sources that are close in temperature as possible. This way we can reduce any in-efficiency as much as possible because, remember that inefficiency is defined as \( \nu = \frac{W}{Q_1}\)  where W is the work done and \(  Q_1 \) is the heat got from the heat source as opposed to the sink(reservoir at the colder temperature). Now one obvious way of getting the best efficiency namely, 1 is to make \( W= Q_1 \) but is this possible? The answer is a resounding no. This leads us to two statements that are the second law of thermodynamics which  are actually equivalent. The second law of thermodynamics says that the efficiency of an engine can never be one and that in fact there is an upper bound that is less than 1.

Second Law of Thermodynamics

1. It is impossible to construct an engine whose sole purpose is to extract heat from a source and convert all of it into work.

2. It is impossible to construct a device that operating in a cycle has the sole purpose of transferring  heat from a cold reservoir to a hot reservoir with no other effect.

As will later be proved these statements are indeed equivalent but they both infer the existence of an amount of heat, \( Q_2 \) which leaks from the engine to the  cold reservoir  so that not all the original heat  \( Q_1\) is converted into work. Therefore the amount of work done can never equal to \(Q_1 \) but must be equal to \( Q_1 -Q_2 \). Plugging this into the equation for efficiency we get that 
\( \nu = 1 - \frac{Q_2}{Q_1}\).

Let us now prove that there exists an upper-bound  \( \nu_c \). This will be proved by contradiction. We shall assume that the upper-bound can be violated and show that this would involve a violation of the second law.

Theorem: There exists an efficient engine,C, with efficiency,\( \nu_c \). This will extract heat ,\( Q_{c_1} \) from the heat source and leak heat \( Q_{c_2} \)to the cold reservoir. This engine C has the property that  \( \nu_c \) is the upper-bound for the efficiency of any system.

Proof
Let us assume there is a more efficient engine,E, than our brilliant engine and let us further assume that they both do the same work difference.  So we have that
\(\frac{W}{Q_1}=\nu > \nu_c \) therefore it must be that \( Q_{c_1}> Q_1\) since we are assuming that \( W=W_c \). Now comes the key idea or trick.

Let us imagine a composite system composite system composed of these two engines, E and C except that the engine that is working at efficiency \( \nu_c\), C, is working backwards (it is a refrigerator) and the hypothetical engine , E,that has more  efficiency  than our upper-bound is doing work that is being put as an input to run the refrigerator.
Recall, that an engine working backwards requires us to put in work so that energy from a cold reservoir can be put into a hotter reservoir. In other words,the engine that is violating our bound on efficiency is running the refrigerator.
To make everything very explicit, E is extracting \( Q_1 \) from a hot reservoir and doing work, W, and leaking \( Q_2\) into the cold reservoir. Engine C is using this same work W, to extract \( Q_{c_2}\) from the cold reservoir and placing into into \( Q_{c_1}\) into the hot reservoir.
This means that the composite system is extracting from the cold reservoir positive  \(Q_{c_1} -W - (Q_1- W)= Q_{c_1} -Q_1  \) and placing the same amount of work into the hot reservoir and the reservoirs are unchanged by the amount of heat added or extracted and stay at the same temperature. But this entails a violation of the second formulation of the second law of thermodynamics. This device  extracts energy from the cold reservoir and put all of it into the hot reservoir with no work done.
      So what has gone wrong? Well, we got that \( Q_{c_1}> Q_1\) which followed from the fact that E was more efficient than C. So this assumption has to be wrong.What we have at this stage is that there is an upper bound to the efficiency that must be less than one. The engine that produces this upper-bound famously goes by the name Carnot's engine and the cycle that the engine goes through in-order to produce this efficiency is called a Carnot cycle.
     Note that we have not mentioned anything about entropy or disorder. We shall see that as a consequence of the formulation of the second law we shall prove the existence of a state function whose change in values from one equilibrium state to another is never negative. This we shall call entropy.
 The second statement of the second law is what Clausius had in mind when he discovered or defined the notion of entropy. It must be emphasized that entropy will then be given the notion of disorder but we can not do that now as we do not have yet the concept of micro-states or atoms as was true in Clausius' time and also by the fact that thermal physics only cares about properties of macro-systems in equilibrium and says nothing about micro-states.

Friday, January 2, 2015

First Law of Thermodynamics

It is interesting to note just how much has been already covered without explicitly mentioning the famous three laws of thermodynamics. It is very tempting to begin with mentioning the three laws and then studying thermal physics, after all supposedly everything banks on them. It is my opinion that this mode of proceeding while it works well for classical mechanics or quantum mechanics does not work well (pedagogically) for thermal physics.  Rather than stating the laws and seeing the consequences as is done in classical mechanics with Newton's laws , it is far better in thermal physics to see how the laws emerge.  Remember our goal is to find a more natural way of formulating the second law that fits well in Thermal physics and does not involve bringing in concepts that are better left in statistical mechanics.

We begin our discussion with an experimental fact that was discovered by Joule.

 If a thermally isolated system is brought from one equilibrium state to another the work necessary to achieve this change is independent of the process used.

Note: Before we argued correctly that the amount of work done was path dependent. But for the special case of adiabatic work it is path independent.

So we are going from one equilibrium state to another and calculating some quantity \( W_{adiabatic}\) that is path independent (does not depend on its history). Recall, this was the special and lovable feature of state functions. Therefore this existence of this experimental fact implies the existence of a state function so that  \( W_{adiabatic} = U_2 - U_1 = \Delta U \). The choice of using the letter U is supposed to be suggestive. Remember the system is thermally isolated and yet we are able to change it from one equilibrium state to another. This must mean that we are changing its internal energy.

But let us suppose that the system is not thermally isolated. Then as pointed out before, the work done on it depends on the path taken. Therefore we can imagine to sorts of scenarios: We have a system A and we done work on it to bring it from one equilibrium state ,a, to another,b, while it is thermally isolated to get  \( \Delta U \) but we can also put it into thermal contact with its surrounding and perform enough work to change it from equilibrium state a, to equilibrium state b. We will not do the same amount of work for both these processes. The difference between these two is heat, Q i,e Q = \( \Delta U - W\). Rearranging, the variables we arrive at what is called the first law of thermodynamics

\begin{equation}
 \Delta U = Q + W  \hspace{10mm} \text{eq.1}
\end{equation}

Note that in contrast to other discussions we have not talked about whether the process is reversible or not ,whether it is quasi-static or not. This is what gives it its generality. Heat therefore is the non-mechanical exchange of energy between the system and the surroundings because of their temperature difference.

There is a subtle difference between what we mean by heat and what we mean by work, or rather what physically heat means as opposed to work. When we have heat put into a system we increase he random motion of the constituent molecules, but when we change the energy of a system by doing work we displace the molecules in an ordered way. From a quantum mechanical perspective, n particles can exist in discrete energy levels \( \epsilon_i \). Let us say that there are \( n_i \) particles on level  \( \epsilon_i \). The total internal energy U = \(  \sum n_i \epsilon_i \). When we do work we change the levels \( \epsilon_i \) with the populations remaining the same, but the populations \( n_i \) can be changed and this is heat.

We can now use the first law of thermodynamics to justify the statement that heat lost is heat gained. This is a crucial assumption or may a obvious assumption used in carnot cycles and engines.
Imagine we a system that we have an adiabatic system so that no heat comes in or goes out and this system and has two subsystems A and B. We can apply the first law to each subsystem separately as follows:
\begin{eqnarray}
 U_f^A - U_i^A &=& Q^A + W^A \\
 U_f^B - U_i^B &=& Q^B + W^B
\end{eqnarray}

We add the two equations to get

\begin{equation}
U_f^A + U_f^B - (U_i^A + U_i^B) = Q^A +Q^B + W^A + W^B
\end{equation}

The term of the left hand side of the above equation is the change in internal energy of the composite system. The first term on the right hand side is the net flow of heat in the composite system and the second term is the work done on the composite system. But remember our assumption was that the system was adiabatic so it must be that  \(Q^A + Q^B = 0\) and therefore \( Q^A = -Q^B\).

Thursday, January 1, 2015

Reversible Processes and Work

Reversible Process 

We are interested in processes that take us from in equilibrium state to another. I  general this will be done irreversibly. But the wonder of state functions is that we can describe or imagine an ideal process that does the job. A simple definition of reversible is that it should be able to return to its original state but this is not all. This process must leave its surroundings unchanged.
Consider the following system: A pendulum and gravitational field. We can apply a force to swing it so we are the environment or the surrounding acting on the system. Imagine pushing the pendulum with an infinitesimal amount of force to move it from one angle to another. We do this to ensure that the pendulum is stays in a state of equilibrium along its path. Ensuring this is what is called a quasi-static process. Now imagine reducing this force so that the pendulum swings back and does work on us and return to its original state. The amount of work it does on us(environment) will be equal to the amount we applied as long as there are no dissipative forces otherwise some work will have to be done to fight against the dissipative forces. Therefore the pendulum has left its surrounding unchanged.
Therefore moving through equilibrium state is a quasi-static process and therefore a reversible process is a quasi-static process where no dissipative forces are present.

Work 
Let us consider, the all to common system involving a gas in cylinder covered by a piston one side. The gas is kept in by a balancing force from the frictionless piston. If the piston is let go slowly enough so that the expansion is quasi-static, the gas will start to expand. All the work done by the gas goes to the environment since there are no dissipative forces. In this simplified and ideal case we can easily calculate the work done by the gas.
\begin{eqnarray}
  dW &=& F dx \\
        &=& P A dx\\
        &=& P dV  \hspace{10mm} \text{reversible processes}
\end{eqnarray}








Therefore we have that:
\begin{eqnarray}
 W = \int_{v_1}^{v_2} P\, dV
\end{eqnarray}

This equation has been derived for reversible process but there are special irreversible processes where it can still be applied because there are no finite pressure drops. So, imagine a cylinder and piston system but this time put small reacting solids that slowly produce gas. This gas does work on the piston which is frictionless. Pressure in cylinder is only infinitesimally smaller than the surrounding and piston is always in mechanical equilibrium and surrounding has pressure  \( P_o\)  so

\begin{equation}
W = \int_{v_1}^{v_2} P\, dV = \int_{v_1}^{v_2} P_o \,dV
\end{equation}

Notice the process is not reversible since the chemical reaction that produces gas is not reversible.
Looking at a P-V diagram we clearly see that the total work done is path dependent as it is the integral under the curve. Therefore dW is and in-exact differential .
Note: The sign 

Example Calculations

1. Suppose we have an ideal gas at T= 300K and isothermally compressed from 1 atm to 10 atm. What is the work done on the gas.

We have that \( \int_{v_1}^{v_2} P \,dv \) but we are not told the volumes but instead told the pressure so obviously we need to change our integration variable. We the use the ideal gas law (PV = nRT) to get the relation between dV and dP . So we have \( dP = - \frac{nRT}{V^2} dV\). So we have that  \(  \int _{P_1}^{P_2} \frac{PV^2}{nRT} \,dP\) we can replace V using the ideal gas law to finally have the integral \(  \int_{P_1}^{P_2} \frac{nRT}{P} dP \) and therefore the formula for work done is \( nRT \ln(\frac{P_2}{P_1}) \). This gives the answer of 57431 J.

2. During and adiabatic irreversible process \( PV^ {\gamma} =c \) all through out the process (not just in this question but in general). Calculate the work done by the gas. \( \gamma \) and c are constants.

The process is reversible so we can use \( \int_{v_1}^{v_2} P \,dV \). We can plug in what P is in terms of and do the integral to get c\( (V_2^{-\gamma+1}-V_1^{-\gamma +1}) \frac{1}{-\gamma +1} \). We have that  \( P_1 V_1 ^{\gamma}=c\) and \( P_2 V_2 ^{\gamma}=c \). We pull back c into the two terms to get  \( W= \frac{P_1V_1 - P_2V_2}{\gamma-1} \)

3. Imagine a gas of n moles at high pressure in a container with volume V and with diathermal walls. The gas is let to leak out slowly through a valve to the surrounding at atmospheric pressure \( P_o\).

This is an irreversible process but it is one of those irreversible processes that is special. The point is that the gas is leaking out very slowly so that there is no finite pressure drop. So we can still use our beloved formula to get \( W= P_o(nv_o -V) \). Here \( v_o \) is the molar volume of the gas outside the container.

The Notion of Equilibrium

Thermodynamics is a perfect science any small inconsistency and the whole structure is destroyed. It therefore pays to be very clear and explicit about each little detail that is to be brought forth to understand anything. For now, we concentrate on what we mean by equilibrium and slightly formalize what we mean by it.
First we curve out a universe for ourselves which consists of everything we care to know about, we call this the system.  Anything that is not contained in it is called the environment or the surrounding. The two spheres of course do not live un-related by are separated by a boundary or a wall.
For a closed system the boundary does not allow for any interaction(no matter exchange) while an open system does allow for matter exchange.

ASIDE: The same kind of taxonomy appears also in quantum mechanics. There for a closed system we have nice unitary evolution wonderfully described by the Schrödinger’s equation and for open systems we have non-unitary evolutions in which classical behaviour emerges. It is often said sloppily that quantum is for the small and classical is for the big, it is more accurate to say quantum is for closed systems and classical is for open systems.

Let us consider a closed system. From experience we know that after a while system reaches a state when no changes occur. In particular, pressure becomes uniform. This state can be labelled by two independent variables (P,V). The amazing thing is that knowing P,V and the masses in the system fixes all other bulk properties that one might what to know about.
We therefore arrive at what we mean by and equilibrium state: this is one in which all bulk physical properties of the system are uniform throughout the system and do not change with time.

I have chosen the pair of variables P,V  but in fact we can choose two any two independent variables e.g for a wire you might choose the tension and the length as your variables. These pairs are called thermodynamic variables or co-ordinates.
It  turns out there are functions which take on unique values as a function of these thermodynamic variables at different equilibrium states. These are called state functions. The concept of a state function is very important because of a very important property namely, the values it takes at different equilibrium states do not depend on history of the system. First of all this is how we shall arrive at state functions but this property is important because in the thermal dynamics we often run into irreversible processes that connect two equilibrium states and we rarely ever have nice neat formulae for these processes but that does not matter because of state functions. All we need to do is describe reversible processes connecting these two equilibrium states (which is nice because reversible processes are easy to describe) and we can understand the physics. This is how we shall arrive at the concept of entropy.

We know we can bring two systems together so that they interact thermally and after a while there will be no further changes to take place in pressure and volume. When no further changes take place we say these two systems are in thermal equilibrium. A transitive property applies to states in thermal equilibrium i.e If A is thermal equilibrium with B and B is in thermal equilibrium with C then A is in thermal equilibrium with C. We can of course apply this argument ad infinitum to describe a whole series of systems that are in thermal equilibrium. This is in essence the zeroth law of thermodynamics.

The amazing thing is that as a consequence of the zeroth law we can prove that all systems in thermal equilibrium share one property, we call this the temperature. The proof which is borrowed from Zemansky's book "Heat and Thermodynamics" will follow momentarily.

We first note a subtle distinction often skirted over namely the difference between thermal equilibrium and thermodynamic equilibrium. Thermal equilibrium does not guarantee thermodynamic equilibrium. In order to have thermodynamic equilibrium we must have mechanical equilibrium i.e all forces are balanced and we must have chemical equilibrium i.e no chemical reactions should be taking place between the two systems.

Now at thermal equilibrium it must be that P,V and T are not independent but are related by some functional equation f(P,V,T)=0. This is called the equation of state. For an ideal gas we have the famous PV -NKT=0.

We now present a proof that given the zeroth law there exits one property that all systems in thermal equilibrium with one another share. This of course we know a head of time is temperature.

Suppose we have to systems A and C with thermal dynamic variables being X and Y for A and X'' and Y'' for C. We know there is and equation of state \begin{equation}
f_{AC} (X,Y,X'',Y'') =0  \hspace{10mm} \text{eq.1}
\end{equation}.
Also let us say there is another system B with variables X' and Y' which is in equilibrium with C so that there is another equation of state
\begin{equation}
f_{BC} (X',Y',X'',Y'') =0 \hspace{10mm} \text{eq.2}
\end{equation}

Now we can solve for Y'' in both equations and set the resulting equations equal to each other
\begin{equation}
g_{AC}(X,Y,X'') = g_{BC}(X',Y',X'') \hspace{10mm} \text{eq.3}
\end{equation}

The zeroth law guarantees that A is in equilibrium with B so that
\begin{equation}
f_{AB} (X,Y,X',Y') =0 \hspace{10mm} \text{eq.4}
\end{equation}
but in particular eq.4 and eq.3 must be in fact equal (remember these are state functions and take on unique values at equilibrium states). This means that X'' is an extraneous variable and can be removed so that at thermal equilibrium we have
\begin{equation}
g_{A}(X,Y) = g_{B}(X',Y') = g_{C}(X'',Y'') = t \hspace{10mm} \text{eq.3}
\end{equation}.
So we have that there exists a function which  parametrizes these sets of co-ordinates and these functions are all equal at thermal equilibrium. This common value is empirically known as the temperature.
 We now see the importance of the zeroth law it guarantees the existence of temperature. I often wondered about the importance of the zeroth law. It is mentioned in the first few pages of thermodynamic textbooks and never makes any appearance as one moves forward. Well, now we know it allows us to talk about thermal equilibrium and once that has happened it is forgotten.
 


Monday, December 29, 2014

Second Law of Thermodynamics and Entropy

       One concept that is often talked about  and rarely understood well is the notion of entropy and the second law of thermodynamics. It is common place to say that the second law states that entropy always increasing or remains constant paired with the statement that entropy is a measure of disorder. This is in fact correct but the problem is that it is stated and talked about carelessly before students have a firm grounding in Statistical Mechanics let alone Quantum Statistical Mechanics.  This is a problem not because anything wrong is being taught but because (as is often the case) very subtle and hard concepts are introduced before one has the right intellectual machinery to fully wrap one's mind about the concept; as a result teachers and professors often resort to very vague statements. Confusion is compounded when students are then given concrete problems from introductory textbooks where  these highfalutin statements can't be readily gleaned and in fact never seem to apply.
     The sad truth is that if we were to properly learn these concepts it is better for the concept of entropy to remain mysterious and simply be talked about as merely a state function. The interpretation of entropy as a measure of disorder requires a proper discussion of statistical mechanics precisely the thing that is not done. In fact, the fact that Entropy always increases can be derived by thinking carefully about Carnot cycles and engines. Recall that Clausius coined the term "entropy" from apparently a greek word, " entropia" which according to wiktionary means " a turning towards".
     The reader should be confused at this point if my point has been made. Why? Well, if entropy is a measure of disorder why did Clausius choose this word when he defined entropy. It seems as though it has nothing to do with disorder. The answer is this; our formulation of the second law only makes sense after accepting Boltzmann's work which came after Clausius' work. This implies there is a formulation of the second law that does not require the definition or formulation of entropy. It is this understanding that takes the back sit in these vague discussions and plays the important role when a lot of introductory thermal physics concepts are problems are being taught.
    It must be stressed that I am not trying to imply that something wrong is being taught, rather I am making a pedagogical point how we should be learning the material. It will be the goal of the next few posts to introduce another way of talking about the Second law of thermal dynamics one that makes no specific reference to entropy as a measure of disorder.

Thursday, December 25, 2014

Analog of Schrodinger Equation for Density Matrices

A basic question we can ask is why bother with density matrices don't wave functions work just fine. The answer is wave functions work just fine as long as we deal with closed systems. The beauty of density matrices is they are better way to understand the dynamics of open systems. This point can be dramatized by noting that there are density matrices that have no wave function analog. An example of this is a completely mixed state i.e zero everywhere except on the diagonal or a scalar times the identity matrix. First we would like to derive the analog of schrodinger's equation for density matrices.

We start with what we know namely, Schrödinger’s equations

\begin{equation}
 i \hbar \frac{\partial |\psi (t)\rangle}{\partial t} = H | \psi (t) \rangle  \hspace{10mm} \text {eq.1}
\end{equation}

We can know imagine the wave function at t=0 and posit the existence of an operator whose sole job is to evolve the operator from time to another. We shall call this operator U.  What this means is that

\begin{equation}
U(dt) | \psi (t) \rangle = | \psi (t +dt) \rangle \hspace{10mm} \text{eq.2}
\end{equation}

Thus  \( \psi (t)\rangle  = U(t) | \psi (0) \rangle \) and we place it in eq.1 to get

\begin{equation} i \hbar \frac{\partial U(t)}{\partial t} | \psi (0) \rangle = H U(t) \rangle | \psi (0) \rangle  \hspace{10mm} \text {eq.3}
\end{equation}

The above equation holds for any initial wave function and at any time so we simply have

\begin{equation}
 i \hbar \frac{\partial U(t)}{\partial t}  = H U(t)  \hspace{10mm} \text {eq.4}
\end{equation}

An analog equation can be got for the adjoint of U. This operator turns out to have the property that \( U U^{\dagger} = I \) .
Now recall,we introduced the density operator as being the outer product of the wave function \( \psi \rangle \) so \( \rho = |\psi \rangle  \langle  \psi | \). We therefore we can write

\begin{equation}
\rho(t) = U \rho_o U^{\dagger} \hspace{10mm} \text{eq.5}
\end{equation}

where  \( \rho_o \) is the initial density matrix at initial time.

We can now take the derivative with respect to time on both sides of eq.5 and multiply by \( i \hbar \) on both sides then using eq.4 whe shall arrive at

\begin{equation}
\frac{d\rho(t)}{dt} = \frac{1}{i \hbar} [H(t), \rho(t)] \hspace{10mm} \text{eq.5}
\end{equation}

When we generalize to open quantum system and additional operator appears in eq.5 usually called the Lindblad operator

\begin{equation}
\frac{d\rho(t)}{dt} = \frac{1}{i \hbar} [H(t), \rho(t)] +\mathcal{L}(\rho) \hspace{10mm} \text{eq.5}
\end{equation}

The additonal operator will look different depending on the system under consideration.




A note on the Hilbert Space


This blog post ties off loose ends.
As has been laid out, there is a direct correspondence between the usual picture of wave functions,Schrödinger’s equations and kets and bras introduced by Paul Dirac.
We have already seen that we interpret what happens in Quantum Mechanics as simply being a collection of Hermitian Operators acting on some vectors. For any experiment what we measure are the eigenvalues and the statistics we collect of experiments are simply the expectation values of these operators. So a question we can ask is where do these vectors "live". From physical intuition we know we need some notion of "dot product" , (we are incidentally using our intuition from the normal Euclidean space). For this dot product the order of operators in the dot product should not matter i.e \( \vec{v}.\vec{w} = \vec{w}.\vec{v} \). For the moment we shall not worry about notation. Also we require that \( \vec{v}.\vec{v} \) be positive definite and only zero when \( \vec{v}=0 \). Lastly, we may require linearity namely \( \vec{u}.( \vec{v} + \vec{w})  = \vec{u}.\vec{v} + \vec{u}.\vec{w}\).
In all this we are assuming all the usual properties that are assumed for distance functions and norms. The vectors  obviously constitute a vector space.  Thus what we really have is an inner product space. Lastly, we also have a metric space with additional property that all Cauchy sequences converge in this vector space.

Note: A metric space is a topological space with a distance function defined upon it. At some point I shall have more to say about this.

Thus we can now say what we mean by a Hilbert space by saying it is an inner product space that is also a complete metric space. Not that this definition does not care what our actual vectors look like. In fact if one goes through the desired properties of our dot product one can easily see that these are satisfied by wave functions with the dot product provided by the integral or by what we usually mean by vectors namely column and row things. In fact the usual n-dimensional real space is also a Hilbert space. In other words there are different kinds of Hilbert spaces; the kind we assume when we write down wave functions obeying a differential equation live in a specific kind of Hilbert space called \(  L^2 \) space. But there is another incarnation of Hilbert space we can use that is provided by bras and kets, as far as I know it has no special name.
These two kinds of Hilbert spaces can roughly be thought of as choosing what sort of basis you want. If you choose a continuous basis you arrive at an \( L^2\) space and generally speaking when we use bras and kets we are most likely working in a discrete basis. This connection shall be explored in greater detail once Representation theory is discussed.

Notation: Dirac introduced  \( | v \rangle \) and  \( \langle w| \) to represent the vectors in our Hilbert spaces. For finite dimensional Hilbert space one can think of these as row and column vectors although technically what we have is an element from a vector space,  \( |v \rangle\) and  \( \langle w| \) an element from the dual vector space (a space of linear functionals). Both these vectors spaces in our case have the same dimensions although in general they need not be.

A stream of posts coming

Over the past year since my last post, a lot of things in a number of different fields have crystallized in my mind. Fields of study like Group Theory, Lie Algebras , Representation Theory, General Relativity, Thermal Physics and more recently Quantum many body physics. These might be trivial insights to most but are something for me. There are two obvious personal projects:

1. A big project I would like to finish is a non trivial document on Differential Geometry and General Relativity. I already have a huge document I prepared but the task of writing everything out in latex scares me.

2.Lie Algebras and Group Theory- There are wonderful similarities between Lie Algebras and Group Theory and fleshing them out explicitly is something I want to do.

 My mental state is not stable, with wild swings between huge mental activity, euphoria and intellectual vitality and other times where it seems my brain is sucked into a black hole accompanied with physical lethargy and utter despair. Goal for this coming year is consistency although I am in no position to make this promise to myself.




Saturday, August 31, 2013

Operators as Matrices: Towards Dirac Notation


 In order to make the jump from the wave function representation to Dirac Notation more obvious and natural let us consider the simplest problem usually discussed at great length in quantum mechanics introductory text book, namely a particle in an infinite box. We shall show that we can write the position and momentum operators as matrices.
To start off, we have a particle in an box with infinite potential at some finite boundary i.e the particle can't escape or tunnel and we are assured that at the boundaries the particle's wave function becomes zero. After going through the calculation(we shall not be done here) we arrive at the fact that the wave functions for the different energy states are given by \(\psi_n=  \sqrt(\frac{2}{L}) \sin (\frac{n\pi x}{L}) \) where L is the length of the box and the energies are  \(E_n =  \frac{n^2 \pi^2 \hbar ^2}{2mL^2} \). Now the Hamiltonian is now merely the kinetic energy operator : \( \frac{p^2}{2m} \) that means that,
\begin{align}
\frac{p^2}{2m} \Psi(t) &= \sum a_n \frac{p^2}{2m} \psi_n \\
                                      &= \sum_n a_n E_n \psi_n \hspace{10mm} \text{eq.1}
\end{align}

Remember that we could represent the wave function as a column vector, the elements of which are the coefficients:

\begin{equation}
\begin{pmatrix} a_1 \\ a_2 \\ a_3 \\ \vdots \end{pmatrix}
\end{equation}

Looking at eq.1 we see that we can represent the kinetic energy operator as
\begin{equation}
\begin{pmatrix}
 E_1 &  0 &  0 & 0 & \ldots \\
 0  &  E_2 & 0 & 0 &  \ldots \\
 0  &   0  &  E_3  & 0 & \ldots \\
\vdots & \vdots & \vdots \\
\end{pmatrix}
\end{equation}

What about the position operator, \( \hat{x}\)?
\begin{align}
 \hat{x}\Psi &= \sum_n a_n \hat{x}\psi_n \\
  \hat{x} \psi_n &= \sum _m x_{nm} \psi_ m  \hspace{10mm} \text{eq.2}
\end{align}
We now use the orthogonality of the \( \psi_m \) to find the matrix elements \( x_{mn} \):
\begin{equation}
 \int \psi_m ^{*} \hat{x} \psi_ n = x_{mn} \hspace{10mm} \text{eq.3}
\end{equation}
We multipled both sides by \(\psi^* \)  in eq.2 and use the fact that  \(\int \psi_n* \psi_m = \delta_{mn} \)
Doing the integral using the eigenfunctions of the particle in  a box mentioned in the introductory second paragraph we get,  \(x_{mn} =  \frac{L}{\pi^2}\frac{4mn}{(m^2-n^2)^2}((-1)^{(m-n) }-1)\). For m=n we get the diagonal elements to be L/2. Remember that m represents the row number and n represents the column number. So when we say elements where m=n we mean the diagonal elements. We then go to our equation make m=n and then find the number to go into the diagonal spots in our matrix. If you want to find the number in the first row second column, make m=1 an n=2 in the equation. We therefore have that our matrix representing the position operator to be:
\begin{equation}
\begin{pmatrix}
\frac{L}{2} & \frac{-16L}{9 \pi^2} & 0 \ldots \\
\frac{-16L}{9 \pi^2}& \frac{L}{2} & \frac{-48L}{25 \pi^2}&\ldots \\
0 & \frac{-48L}{25 \pi^2} & \frac{L}{2}& \ldots \\
\vdots & \vdots &  \vdots  & \ldots
\end{pmatrix}
\end{equation}

Notice in eq.3 we sandwiched the operator between the eigenfunction and its complex conjugate in order to find the matrix element.
The same can be done with the momentum operator \hat{p}or represented concretely as \( \frac{-i\hbar d}{dx} \). We have :
\begin{align}
 \hat{p}\Psi &= \sum_n a_n \hat{p} \psi_n \\
   \hat{p}\psi_n &= \sum_m  p_{nm} \psi_ m \\
\end{align}
Using the orthogonality of the \( \psi_n \)  we get another integral to do namely, \(\int \psi_m ^{*} \hat{p} \psi_ n \). This time the integrals gives us \( p_{mn} = \frac{\hbar}{iL} \frac{2mn}{m^2-n^2} (1 - (-1)^{m-n})\). So the matrix representing our momentum operator is:
\begin{equation}
\begin{pmatrix}
 0 & \frac{8i\hbar}{3L} & 0 \ldots \\
\frac{-8i\hbar}{3L}& 0 & \frac{24 i\hbar}{5L}&\ldots \\
0 & \frac{-24 i \hbar}{5L} & 0 & \ldots \\
\vdots & \vdots &  \vdots  & \ldots
\end{pmatrix}
\end{equation}

So as we can see  our matrices are really infinite and they act on our infinite dimensional column.
We have now laid the ground for Dirac notation. We have seen that in this simple problem there is a representation which can be arrived at where we think of ourselves as living in some sort of vector space where our operators(matrices) and our vectors can be infinite dimensional.
<script type="text/javascript">MathJax.Hub.Queue(["Typeset",MathJax.Hub]);</script>
 

Thursday, August 29, 2013

Transition from wave functions to Dirac notation

We begin as one would expect with Schrodinger's equation in one dimension; which is
\( -\frac{\hbar}{2m}\frac{d^2 \psi }{dx^2} + V \psi = \frac{i \hbar \partial \psi}{\partial t} \) . If we solve the solve this using separation of variable we get a term that is only time dependent and the other that is only spatially dependent. The steps to this abound everywhere so I shall not go into it (It would not be in the spirit of this blog).  But we could introduce what is called the Hamiltonian operator and rewrite the time dependent Schrodinger's equation as follows:
\begin{equation}
H\psi = i \hbar  \frac{\partial \psi}{\partial t}
\end{equation}
where H =  \(-\frac{-\hbar}{2m}\frac{ d^2 }{dx^2} + V  \).
The spatially dependent part would look like this; (If we write in terms of our newly defined Hamilton, H)
\begin{equation}
H\psi = E \psi
\end{equation}
Notice that E(energy) is just a number and this looks like some eigenvalue problem which in fact it is. So if we have a system with different energy levels , these would correspond to different eigenfunctions.  So in that case our eigenvalue problem would look like this:
\begin{equation}
H\psi_n = E_n \psi_n
\end{equation}
The next statement is very important and will be stated without proof. The \(\psi_n \) live in a vector space that physicist refer to as the Hilbert Space, not only do they live in it but also can form a basis set for it (after of course they have been normalized). In other words if we have some solution to the schrodinger's equation, we can express it as a linear combination of these  \(\psi_n \). That is to say that:
\begin{equation}
 \Phi  = \sum_n c_n \psi_n
\end{equation}
where \( \Phi(t) \) is a solution to Schrodinger's equation.
It is here that the student may think his or herself, " well that's neat, what is the next topic?"  Pose for a while and contemplate this, we have a infinite dimensional vector space, spanned by these \( \psi_n \). A little but deep thought could occur to one: why represent the vector by writing the tedious sum expressed in the above equation? Could we not just use the \(c_n \)  place them in a column vector (which in this case would be infinite dimensional) and then our operators would be matrices rather than differential operators. After all matrices act upon column vectors.  This is the key insight to Dirac notation. Our quantum state will no longer be written by some function but instead it will be some column vector, with the elements in it the \(c_n\).
In other words or symbols:
\begin{equation}
 \Phi =  \begin{pmatrix}  c_1\\ c_2 \\ c_3 \\ \vdots \end{pmatrix}
\end{equation}

NOTE: If we are in an infinite dimensional space we have to worry about convergence but at the most we live in a textbook world everything is well behaved. The next step is to see the transition from differential operators to matrices explicitly.

Sunday, August 25, 2013

A late realization.(Dirac Notation)

Since I would like to have some sort of structure to the posts on this blog. I realized that I jumped into Density matrix theory without ever talking about Dirac Notation which was heavily used. Since I would like this blog to be completely self contained to the curious reader, I shall stop , back up and slowly introduced Dirac notation and may be in the process introduce Quantum Mechanics.

Density Matrix Theory: An Introduction III

One of the main reason from introducing the density matrix is that it allows to describe statistical mixture as opposed to merely pure states. We therefore have to make some comments as to how we generally extract out information about a system from quantum mechanics. We know from introductory quantum mechanics that the wave function has all the information we could want or that quantum mechanics can provide.
Let's for a moment talk more generally about what is knowable in quantum mechanics. We have all heard about Heisenberg's uncertainty principle, namely that we can't know position and momentum exactly simultaneously. One often hears this in public discussions, but a more precise statement and in fact a more correct statement is that we can't know the position in a certain direction and the momentum in that same direction simultaneously to arbitrary accuracy. The mathematical version of this statement would be  that \([ \hat{q_i},\hat{p}_j ] = i\hbar \delta_{ij}\) . Here \(q_i\) and \(p_i\) are generalized co-ordinates and generalized momentum respectively. So the more familiar version would be \( [\hat{x}, \hat{p_x}] = i\hbar \).
The underlying reason is that the position operator in the \(\hat{x}\) does not commute with the momentum operator in the same position. This implies that I can indeed now the position exactly in the \(\hat{y}\) and the momentum for example in the \(\hat{x}\) .
Now to jump to more general statements, we may proceed as follows. If we have a set T=\( ( Q_1,Q_2,Q_3 ...Q_n) \)  consisting of operators any two of which commute with each other  and a correponding set U = \((q_1, q_2,...q_n)\)  of eigenvalues for each operator ( we assume that this sets are the largest possible size they can be so the ket  \(| q_1, q_2,q_3,...q_n \rangle \) describes our system), then these two sets represent the "maximum knowledge" of the system that we can have. These states of "maximum knowledge"  are what we called the pure states earlier.
One must keep in mind that these two sets may not be unique and in fact  are rarely  reproducible in experiments. Instead what we usually have is that we know with a certain probability  \( W_n \) that our system is in a pure state \( \psi_n \). We therefore deal with statistical mixtures more than we do with pure states.

NOTE: Do not confuse a superposition of states with a statistical mixture. With a superposition there is a phase relationship between the quantum states, we can thus write down a pure state wave function describing a quantum superposition of these two states. On the other hand, statistical states  \( \psi_n \) have no phase relationship. To put it another way, if I have a state and I can't tell in principle whether it is spin up or spin down until I measure then I have a superposition but if I have a number of states that are either spin up or spin down or superposition of the two (so there is probability of me picking up a state with either spin or spin down or superposition of the two) then I have a mixed state.

In quantum mechanics we connect with the "real world" by calculating expectation values  of some Operator. So for example, we continuously prepare a state in some specific configuration and keep measuring its energy. At the end we shall have a probability of having the state with energy \(E_n\)  in the corresponding eigenfunction  \( \psi_n\). We can calculate the expectation value for a pure state (in braket notation) like this:
\begin{equation}
\langle O \rangle = \langle \psi| O | \psi \rangle
\end{equation}
and for mixed states
\begin{equation}
\langle O \rangle = \sum _n \langle \psi_n| O | \psi_n \rangle
\end{equation}

Density Matrix
We can now describe our density matrix of our statistical mixture in the following manner:
\begin{equation}
 \rho = \sum_n W_n |\psi_n \rangle \langle \psi_n|
\end{equation}
Here \( W_n \) are the statistical weights and\( |psi_n \rangle\) are the independently prepared states. These  independently prepared states are not necessarily orthonormal  so we could re-write in terms of states that are so that :
\begin{equation}
\psi_n = \sum_m a_m^n \phi_m
\end{equation}

We are now in a position to re-write our density matrix in the following manner:
\begin{equation}
 \rho = \sum_{nmk} W_n a_m^{n} a_k^{n*} |\phi_m \rangle \langle \phi_k|
\end{equation}

Thus we are able to put matrix in density matrix by find how to get its matrix elements. Using the orthogonality condition of the \( \phi_n \) state we observe that:
\begin{equation}
\rho_{ij} = \langle \phi_i |\rho | \phi_n \rangle = \sum _n W_n a_j ^{n}a_j^{n*}
\end{equation}

Introducing these orthonormal states allows us to arrive at a simple expression or the expectation value for an operator of a mixed state. We proceed as follows:
\begin{align}
 \langle O \rangle &= \sum _n \langle \psi_n| O | \psi_n \rangle \\
                         &= \sum_{mk} \sum_{n} W_n a_m^{n} a_k^{n*} \langle \phi_k |O | \phi_m \rangle \\
                           & = \sum_{n}\langle \phi_m|\rho | \phi_k \rangle  \langle \phi_k |O | \phi_m \rangle  \\
                            &= tr (\rho O)
\end{align}

Look at the power of the density matrix. If I know the density matrix (remember something I can define for pure or statistical mixtures), I can arrive at the very thing I need to connect with  the "real " world namely the expectation of the operator or observable in question. This something that I can't easily get if I stick with the wave function and have a statistical mixture in my lab.

Saturday, August 24, 2013

Density Matrix Theory: An introduction II

We have now introduced the polarization vector and how to calculate its components. To see the relationship between it and the density we shall move as follows:
Remember that if we have  a pure state then we can assign a state vector to our system but if we have a mixture then we instead have to consider a statistical mixture. We can thus write the density matrix for a statistical mixture in the following manner:\chi
\begin{equation} \label{eq.1}
\rho = \sum_k W_k |\chi \rangle  \langle \chi | \hspace{10mm}  \text{eq.1}
\end{equation}
where \(W_k \) is the statistical weight. From this I hope one can see that a statistical mixture is a generalization of a pure state. Let's multiply the above equation by the pauli matrix \( \sigma_i \) to get:

\begin{equation} \label{eq.2}
 \rho \sigma_i =  \sum_k W_k |\chi \rangle  \langle \chi |  \sigma_i   \hspace{10mm}  \text{eq.2}
\end{equation}
We shall now take the trace of each term i.e the \(tr (\rho \sigma_i \) to get:
\begin{equation} \label{eq.3}
tr ( \rho \sigma_i ) = \sum_k W_k \langle \chi |  \sigma_i | \chi \rangle = P_i   \hspace{10mm}  \text{eq.3}
 \end{equation}

Now, the pauli matrices in combination with the identity matrix form a basis for a 4 dimensional space. Thus if we stick with two level systems, then we can write our density matrix as:
\begin{equation}
 \rho = a_o I + \sum_i a_i \sigma_i
\end{equation}
 We can the use the following easily derivable facts to find the \( a_o , a_i  \)constants:
 1) pauli spin algebra- \(\sigma_i \sigma_j = i \sum_i \epsilon_ {ijk} \sigma_k  + \delta_{ij}I\),
 2) the \(tr(\sigma_i) =0 \)
 3)\( tr (\rho ^2 )= \frac{(1+ P^2)}{2}\)

Using  2) and 3) in combination tells us that \(a_o \) is actually 1/2. Multiply our density matrix again by  \(sigma_i \)  and taking the trace and then using 1) and 2) gives us that
\begin{eqnarray}
 tr ( \rho \sigma_ i) &=& 2 \sum_i b_i \delta_{ij} \\
                              &=& 2 b_i
\end{eqnarray}
Now referring back to eq.3  we see that  \( b_i  = \frac{P_i}{2} \)
But now let's take a close look at eq.1 in conjuction with fact 3). In general eq.1 will be less than 1 since the statistical weights themselves are less than one. Which means that in general  \( \rho^2 \) for  a statistical mixture will be less than one while if we have a pure state then \(  \rho^2\) is equal to one. Now we finally have a mathematical way of distinguishing between a density matrix for a pure state and a density matrix for a statistical mixture.

In our discussion of the density matrix we considered states \( | \chi \rangle \) . These states were not necessarily orthonormal. But we could re -write the density matrix in terms of orthornormal states. That is to say we can first re-write the \( | \chi \rangle \) like this:
\begin{equation}
 |\chi  \rangle  =  \sum _n a_n | \psi \rangle
\end{equation}

So if we then re-write the density matrix (which is in general represent a mixture) we have that:
\begin{equation}
 \rho = \sum_k W_k a_k a_k^{*} | \psi \rangle \langle \psi |
\end{equation}

This last equation gives us a hint of a much more general formulation for the density matrix with the eventual goal of considering open quantum systems and the density matrix that may describe them.

Friday, August 23, 2013

Density matrix theory: An Introduction, Preliminary Considerations

A very powerful and useful instrument that will be discussed is the density matrix. Consider the outer product of a ket and a bra. i.e \( |\psi \rangle \langle \psi | \). This might look like a needless complication but as we shall see this entity tells us quite a bit of information.
Before we start with the wonders we shall start with the Stern Gerlach experiment. A beam of electron is sent into a magnetic field pointing in the \(\hat{z} \) direction. What was expected was that electron would be bent or deflected in all directions instead what was seen was that either the electrons were deflected up or down. This was the first indication of what physicists now call spin. I shall not go on about this experiment since it is one that is rehashed in almost every modern physics textbook or quantum mechanics text book. Instead what shall is consider the following scenario:

Supposing a beam of light is let go from a source and it is able to move through the Stern Gerlach instrunmentation by turning the instrument in some direction, then we shall call it a pure state. Note that we have not specified a direction, we have merely said a direction or orientation can be found in which the beam entirely goes through.
This is important because if  particles  in the  beam  are in a superposition of spin up and spin down, as long as all the particles are in this state then this beam is considered to be in a pure state. If this is the case we can then describe the system by a single state vector.

If one the other had we have beam of particles in one pure state and another beam of particles in another pure state, then this ensemble is considered to be in a mixed state. In order for this to occur the beams have to be prepared independently and by this we mean that no phase relation exists between these two beams.

Now we move to introduce the polarization vector which shall be defined as follows:
$$ P_i = \langle \sigma_i \rangle =\langle \chi |\sigma_i | \chi \rangle  $$
If we are talking about a two level system then the sigma are the pauli spin matrices. Before we continue, let's get an inkling as to why it is called the polarization vector. If we consider the general two dimensional state vector \( | \chi \rangle = \cos \theta + e^{\delta}\sin \theta \)  and then calculate \(P_i\), one gets the following column vector \(  (\cos \delta \sin \theta, \sin \delta \sin \theta, \cos \theta)^{T} \) This should remind one of spherical co-ordinates. So we picture the bloch sphere the polarization is telling us the "composition" of our state vector. Hence how it is polarized, in very much the analogous way we might ask how the electric field for example is polarized. In fact there is an analogy to be made between filter EM waves and spins the Stern-Gerlach experiment.

The definition above should just remind one of a pure beam, well can we generalize it for mixtures and as a point in fact it can, namely:

$$ P_i = \sum _a W_a \langle \chi_a |\sigma_i | \chi _a\rangle $$

Here \( W_a \) are simply the statistical weights namely the proportion of particles in state \( \chi_a \)

The next step is to connect this with density matrix.