Wednesday, May 29, 2013

Heap Data Structure

Heap is an efficient data structure used for managing data during the execution of algorithms like heapsort and priority queue. It supports insert and extract min (or max) operations efficiently. Heap achieves this by maintaining a partial ordering on the data set. This partial ordering is weaker than the sorted order but stronger than random or no order. Heap is a binary tree in which key on each node dominates key of its children.

There are two variations of heap, min-heaps, and max-heaps(as shown below).


                                    10  <  45,75                                      2001 > 1045,1175
                                    45  <  70,55                                      1045 > 70,55

Property of Heap

  1. The tree is completely filled on all levels except the lowest level which gets filled from left up to a point. So, it can be viewed as the nearly complete binary tree.
  2. Both heaps(min and max) satisfy a heap property. In max-heap, the key of each node is larger than the key of its children nodes. A min-heap is organized the opposite way.
  3. Smallest element in min-heap is at the root. Largest element in max-heap is at the root.
  4. Heap is NOT a binary search tree (it is a binary tree though ). So binary search doesn't work on the heap.

Heap Implementation

  1. Binary Tree: The obvious implementation of heap is binary tree (evident from above diagram). Each key will get stored in node with two pointers to its children. Height of heap is the height of root node; and height of a node is number of edges on the longest downward path from node to the leaf. So height of above trees are 3 ( lg8 ).
  2. Array: Nearly complete binary tree in case of heap enables it to represent it without any pointers. Data can be stored as an array of keys, and use the position of the keys to implicitly satisfy the role of the pointers. 
         
Node(i)left = 2* Node(i)  
Node(i)Right = 2*Node(i)+1


       As shown above, root of the tree is stored at the first position of the array, and its left and the right children in the second and third positions, respectively. Assumed that array starts with index 1 to make calculations simpler. So left child of i sits in position and right child in 2i+1. This way you can easily move around without pointers.        

           int first_child(int i){
               return 2*i;
           }
           
           int parent(int i){
              if(i == 1) return -1
              else return (int)i/2;
           } 

 Array representation of heap saves memory but is less flexible than tree implementation. Updating nodes in array implementation is not efficient. It helps to implement heap as array as its fully packed till last level. But it will not be so efficient for normal binary trees. 

Heap Construction

The fundamental point which needs to be taken care of is; the heap property should be retained during each new insertion. Let's take the case of the above min-heap. Now assume that you want to add one more item to the heap with key 5.

Step 1: Add 5 to the next slot. 
The new element will get added to the next available slot in the array. This  ensures the desired balanced shape of the heap-labeled tree, but heap property might no longer be valid.



Step 2: Swap Node 5 and 70 to satisfy the min-heap property. 
After swap min-heap property test passes for node 5. But min-heap property fails for the parent node of 5.



Step 3: Swap node 45 with 5.
Now node with key 5 is satisfied as its children have key values > 5.



Step 4: Swap root node with its left child.
Bingo. Heap is perfect now.




Insertion Complexity: Swap process takes constant time at each level. So each insertion takes at most O(log n) time.

Creation Complexity: Each insertion takes O(log n) time; then creation (or insertion of n nodes will take O(n log n) time.


Extracting Minimum from min-heap

Minimum of a min-heap sits at first position; so finding first minimum value is easy. Now the tricky part is how to find next minimum element. 

To find the next minimum; let's remove the first element. After this to make sure that binary tree is maintained, let's copy the last element to the first position. This is required to ensure that the tree is filled at all levels except the lowest level. But after this, the heap property might go for a toss as now a bigger element may sit at the root. Now the heap should re-arrange itself so that the min-heap property satisfies. 

Heapify:  This is a process in which a dissatisfied element bubbles-down to appropriate position. To achieve this the dominant child should be swapped with the root to push down the problem to the next level. This should be continued further till the last level. 
The heapify operation for each element takes lgn steps (or height of tree) so complexity is O(log n) time. 

Java Implementation: link

Heapsort

Exchanging the minimum element with the last element and calling heapify repeatedly gives an O( n logn) sorting algorithm, named as Heapsort. Worst case complexity is O( n logn ) time, which is best that can be expected from any sorting algorithm. It is an in-place sorting i.e. uses no extra memory 




Saturday, May 4, 2013

Class Loading and Unloading in JVM

Class Loader is one of the major component (sub system) of the Java Virtual Machine(JVM). This posts talks about loading and unloading of a class (or interface) into virtual machine. JVM specification refers to classes and interfaces as type.

A Java program is composed of many individual class files, unlike C/C++ which has single executable file. Each of these individual class files correspond to a single Java class and it gets loaded the first time you create an object from the class or the first time you access a static component (block, field or method) of the class.

A class loader's basic objective is to service a request for a class. JVM needs a class (ofcourse for instantiating it), so it asks the class loader by passing full name of the class. And in return class loader gives back a Class object representing the class(more detail here). Below diagram shows sequence of steps:



Class Loading

In the JVM architecture tutorial (link), I have discussed class loader subsystem briefly. JVM has a flexible class loader architecture that enables a Java application to load classes in custom ways. Each JVM has at 
least two class loaders. Let's cover both of them:
  • Bootstrap class loader : Part of JVM implementation and it is used to load the Java API classes only. It "bootstraps" the JVM and is also known as primordial, system or default class loader.
  • User-defined class loaders : There could be multiple user-defined class loaders. Class loaders are like normal Java classes, so application can install user-defined class loaders to load classes in custom way. You can dynamically extend Java application at run time with the help of user-defined class loaders.
JVM keeps track of which class loader loaded a given class. Classes can only see other classes loaded by the same class loader (for security concerns). Java architecture achieves this by maintaining name-space inside a Java application. Each class loader in a running Java application has its own name-space and class inside one class loader can't access a class from another class loader unless the application explicitly permits this access. This name-space restriction means :
  • You can load only ONE class named as say Fruit in a given name-space. 
  • But you can create multiple name-space by creating multiple class loaders in the same application. So, you can load three Fruit classes in three different class loaders. And all these Fruit classes will be unaware of the presence of rest two.
Now, obvious question is, how these multiple class loaders in the same application load classes ?
Class loaders use Parent-delegation technique to load classes. Each class loader except bootstrap class loader has a parent. Class loaders asks its parent to load a particular class. This delegation continues all the way to the bootstrap class loader, which is the last class loader in the chain. If the parent class loader can load a type, the class loader returns that type. Otherwise current class loader attempts to load the class itself. This approach ensures that you can't load your own String class, because request to load a new String class will always lead to the System/bootstrap class loader which is responsible for loading Java API classes. Also classes from Java API get loaded only when there is a request for one. 

What if, I create my own class in Java.util package ?
Java gives special privileges to classes in the same package. So does it mean my class say Jerk inside Java.util package will get special access and security will get compromised ?  
Java only grants this special access to class in the same package which gets loaded by the same class loader. So your Jerk class which would have got loaded by an user-class loader will NOT get access to java.lang classes of Java API. So code found on class path by the class path class loader can't gain access to package-visible members of the Java API. 

Class Unloading

Lifetime of a class is similar to the lifetime of an object. As JVM may garbage collects the objects after they are no longer referenced by the program. Similarly virtual machine can optionally unload the classes after they are no longer referenced by the program. 

Java program can be dynamically extended at run time by loading new types through user-defined class loaders. All these loaded types occupy space in the method area. So just like normal heaps the memory footprints of method areas grows and hence types can be freed as well by unloading if they are no longer needed. If the application has no references to a given type, then the type can be unloaded or garbage collected (like heap memory).

Until Java SE 7; classes (and Constant pool) were stored in Permanent Generation (or PermGen).  Tuning PermGen size was complicated so Java SE 8 has completely removed PermGen and now classes (Java HotSpot VM representation of class) are moved into native or heap memory, link.

Types (Class/Interface) loaded through the bootstrap loader will always be reachable and will never be unloaded. Only types which are loaded by user-defined class loaders can become unreachable and hence can be unloaded by JVM. A class instance can be reachable in following case:
  • Class instance will be reachable if the application holds an explicit reference to the instance. 
  • Class instance will be reachable if there is a reachable object on the heap whose type data in the method area refers to the Class instance. 

References:
http://www.artima.com/insidejvm/ed2/jvm5.html 

Related Articles:
how-garbage-collection-works
jvm-architecture
life-cycle-of-object-in-java

---

calling it a post; do give your feedback!!!

Wednesday, May 1, 2013

Life Cycle of an Object in Java

In Java, objects play the pivotal role; so understanding its instantiation, and how and when it gets garbage collected is important. This post covers the life cycle of objects inside Java Virtual Machine. But keep in mind that, successful creation of the object is possible only if class is already loaded in JVM (by class loader).


Object Creation 

Class instantiation activity gives life to an object and thereafter you can call its methods to manipulate states/attributes. Object creation is done by Java Virtual Machine(after class has been loaded). When JVM creates a new instance of a class, it allocates memory on the heap to hold the object's instance. Memory is allocated for all variables/attributes declared in object's class and in all its superclasses (including hidden attributes). Once the virtual machine has allocated memory for the new object, the machine initializes the instance variable to default initial values. After this VM gives proper initial value depending on how the object gets created:

   1.  new keyword
        Person p = new Person();
        String s = new String("life");
            This is one of the most frequently used approach to create an instance of a class. 

       2.  newInstance() on a Class / Reflection
            Class class = Class.forName("fully.qualified.class.name");
            Object obj = class.newInstance();
            Person p = (Person)obj;
             

       3. Cloning         
           Person pClone = (Person)p.clone(); 

           Make sure that, Person class implements Cloneable interface and also provides the definition of the clone() method. 

       4. Deserialization      
           FileInputStream fis = new FileInputStream("person.ser");
           ObjectInputStream ois = new ObjectInputStream(fis);
           Person per =(Person)ois.readObject(); 
          
           Deserialization is a technique to convert the encoded byte stream into the corresponding object as 
           shown above. Refer to the detailed tutorial.

       5. Implicit Instantiation
           All above techniques explicitly creates an object of a class. Apart from above, there are multiple implicit ways of object creation as well :
              
           1. Class loading : For every type that a JVM loads, it implicitly creates a new Class object to                  represent that type.
           2. String Literals : When JVM loads a class which has String literals (i.e. String PI = "3.14"), 
               it might create a new String object. 
           3. Expression Evaluation : Process of evaluation an expression like String concatenation
               can also create new objects. 
                   
     So above explicit/implicit object instantiation techniques give proper initial value to the object. JVM uses one of the three techniques to do this task, depending on how the object is being created.

    If the object is being created because of a clone() invocation, the virtual machine copies the values of the instance variables of the object being cloned into the new object. If the object is deserialized via a readObject() invocation on an ObjectInputStream, the virtual machine initializes non-transient instance object variables from values read from the input stream. Otherwise, the virtual machine invokes an instance initialization method on the object. The instance initialization method initializes the object's instance variables to their proper initial value. 
    The Java compiler generates at least one instance initialization method for every class it compiles. In the  Java class file, the instance initialization method is called "<init>". For each constructor in the source code of a class, the Java compiler generates one <init>() method.  

    Object in action

    Once object is created successfully, its ready for the real action.
    obj.haveFun();

    Garbage Collection of Object / End of life

    Application can allocate memory for object via explicit and implicit ways described under object creation, but they can not explicitly free that memory. To signal to JVM that an object is ready for garbage collection, object needs to get unrefrenced in either of below ways :

    1. Explicit unreferencing
        person = null;  //person is an instance of Person   

    2. Object goes out of scope
        An Object created inside a method goes out of scope once method returns to the calling method.
        So in below case once method() is over, p implicitly becomes null. 

        public void method(){
           Person p = new Person();
           //method action
        } 

       After unrefrencing, VM can reclaim that memory. Virtual machine can decide when to garbage collect unrefrenced object. But specification doesn't guarantee any predictable behavior. It is up-to the VM to decide when to reclaim memory from unrefrenced object or it might NOT reclaim memory at all.
    If the class declares a finalizer (i.e. public void finalize() method ) then Garbage Collector will execute the finalize() method on the instance of the class before it frees the memory space occupied by that instance. So clearly, exact time of garbage collection is unpredictable.


    Sunday, April 28, 2013

    Latch API in Java : CountDownLatch

    Literally, latch means a device for keeping a door or gate closed.  Its meaning is analogous to a gate in Java as well. So if latch (gate) is open, everyone can pass through it but when it's shut, no one is allowed to cross over. With this as background let's go in detail:

    It's one of the advanced threading/concurrency concepts of Java. Java provides a Latch API named as  CountDownLatch which got introduced in Java 5 ( java.util.concurrent package). So now on; I will refer Latch and CountDownLatch interchangeably to refer to the same thing. Latch is a synchronizer that can delay the progress of threads until it reaches its terminal state. [A synchronizer is any object that coordinates the control flow of threads based on its state] So it is used to synchronize one or more tasks by forcing them to wait for the completion of a set of operation being performed by other tasks.

    CountDownLatch in action

    Let me start directly with a simple example to stress on the fundamentals of this API. Below sample program has two tasks (as taskone() and tasktwo()) represented as methods. And I want to make sure that taskone() should get completed before tasktwo(). 

    import java.util.concurrent.CountDownLatch;
    import java.util.concurrent.TimeUnit;
    
    public class CountDownLatchTest {
     static CountDownLatch latch;
    
     CountDownLatchTest(final int count) {
      latch = new CountDownLatch(count);
     }
    
     public void firstTask() {
      Runnable s1 = new Runnable() {
       public void run() {
        try {
             System.out.println("waiting....");
             TimeUnit.SECONDS.sleep(5);
        } catch (InterruptedException e) {
             e.printStackTrace();
        }
        // finish first activity before last line
        latch.countDown();
       }
      };
      Thread t = new Thread(s1);
      t.start();
     }
    
     public void secondTask() {
      Runnable s2 = new Runnable() {
       public void run() {
        try {
             System.out.println("wait....");
             latch.await();
             // perform task here
             System.out.println("after wait.... done");
        } catch (InterruptedException e1) {
             e1.printStackTrace();
        }
       }
      };
      Thread t = new Thread(s2);
      t.start();
     }
    
     public static void main(String[] args) throws InterruptedException {
      CountDownLatchTest cdlt = new CountDownLatchTest(1);
      cdlt.secondTask();
      TimeUnit.SECONDS.sleep(5);
      cdlt.firstTask();
     }
    }
    

     Output:
     wait....
    waiting....
    after wait.... done

    Above example shows CountDownLatch attribute getting initialized to a value of 1 through constructor. Two tasks in above example are synchronized though CountDownLatch. Task2 i.e. secondTask() should wait for completion of Task1 i.e. firstTask(). Run above example and notice the sequence in which output appears on the console. 

    Please note few important points

    1. Any task that calls await() on the object will block until the count reaches zero or it's interrupted by another thread. secondTask() gets blocked after call to await(); evident from output.
    2. Call countDown() on the object to reduce the count. The task that call countDown() are not blocked.  This cal signals the end of the task. This method need to be called at the end of the task.
    3. As soon as count reaches zero; threads awaiting starts running. 
    4. The value of count which is passed during creation of latch object is very important. It should be same as the number of task which needs to be finished first. If count is 5 then first task should be called five times to make sure that count has reduced to 0.
    5. You can also use wait and notify mechanism of Java to achieve the same behavior but code will become quite complicated. 
    6. One of the disadvantage of CountDownLatch is that its not reusable once count reaches to zero. But Java provides another concurrency API called CyclicBarrier for such cases. 

    Usage of CountDownLatch

    1. Use this when your current executing thread/main thread needs to wait for the completion of other dependent activities. 
    2. Ensure that a service doesn't start until other services on which it depends have not completed.
    3. In a multi-player game like RoadRash; wait for all players to get ready to start the race. 

    ---
    do post your comments/questions !!!

    Sunday, April 21, 2013

    Semaphore in Java

    Let's start with below problems; before starting on Semaphore.
    1. Implement a Database connection pool which will block if no more connection is available instead of failing and handover connection as soon as it's available
    2. Implement a thread pool
    3. Create a blocking bounded collection
    4. Limiting number of http connection to an external site
    Note : Bounded puts an upper limit on the size of collection. 
    And block(ing) here means that, wait for the collection to become non-empty when retrieving an element, and wait for space to become available when storing an element. 
    (One of the) Solution to all above problems is, Semaphore!
    Recall that, in normal monitor lock (implicit lock), only one thread can access a resource. So, if you want more than one thread to access a resource at the same time, Semaphore comes into play. Semaphores are used to restrict number of threads which can access a resource or perform a given action at the same time. To achieve this, semaphore maintains a counter which keeps track of the number of resources available.

    When a thread requests access to resource, semaphore checks the variable count, and if it's less than total count, then grants access and subsequently reduces the available count. If count is equal to maximum allowed count then it asks thread to wait. If resource count is one (0/1 or on/off) then its called as Binary Semaphore, otherwise its called as counting semaphore.

    Semaphore got introduced into standard Java library in version 5.

    Limit Number of Http Connections

    Below example show how can you limit the number of http connection in your system.

    import java.io.IOException;
    import java.net.URL;
    import java.net.URLConnection;
    import java.util.concurrent.Semaphore;
    
    /**
     * Limit Maximum number of URL connection allowed through Counting Semaphore
     * @author Sid
     *
     */
    class UrlConnectionManager {
     private final Semaphore semaphore;
     private final int DEFAULT_ALLOWEED = 10;
    
     UrlConnectionManager(int maxConcurrentRequests) {
      semaphore = new Semaphore(maxConcurrentRequests);
     }
     
     UrlConnectionManager() {
      semaphore = new Semaphore(DEFAULT_ALLOWEED);
     }
    
     public URLConnection acquire(URL url) throws InterruptedException,
       IOException {
      semaphore.acquire();
      return url.openConnection();
    
     }
    
     public void release(URLConnection conn) {
      try {
       // clean up activity if required
      } finally {
       semaphore.release();
      }
     }
    }
    

    In above example; semphore maintains a set of permits or pass. Method, semaphore.acquire(), blocks if necessary until permit is available and then takes it. And each semaphore.release() adds/returns a permit. Semaphore is like a gatekeeper which keeps track of number of visitors allowed in a building.

    Blocking Bounded LinkedList

    You can also use semaphore to turn any collection into blocking bounded collection. Let's see how to do the same for LinkedList. Assume that, only add and remove operations are allowed. Below class provides main method to test it. Note that, at most 5 add operations are allowed.

    import java.util.Collections;
    import java.util.LinkedList;
    import java.util.List;
    import java.util.concurrent.Semaphore;
    import java.util.concurrent.TimeUnit;
    
    /**
     * Minimilistic Bounded List  
     * @author Sid
     *
     */
    public class BoundedLinkedList<T> {
     private List<T> list;
     private Semaphore semaphore;
    
     public BoundedLinkedList(final int bound) {
      list = Collections.synchronizedList(new LinkedList<T>());
      semaphore = new Semaphore(bound);
     }
    
     public void add(T o) throws InterruptedException {
      semaphore.acquire();
      list.add(o);
      System.out.println(" Added element :"+ o);
     }
     
     public void remove(T o) {
      if (list.contains(o)) {
       list.remove(o);
       semaphore.release();
       System.out.println(" Removed element :"+ o);
      }
     }
     
     /**
      * Test Method
      * Adds 6 elements; then waits for 5 seconds; removes one element
      */
     public static void main(String[] args){
         final BoundedLinkedList<Integer> bll = new BoundedLinkedList<>(5);
      new Thread(new Runnable(){
       public void run(){
        for(int i = 1; i <= 6; i++)
         try {
          bll.add(i);
         } catch (InterruptedException e) {
          e.printStackTrace();
         }
       }
      }).start();
      
      try {
       TimeUnit.SECONDS.sleep(5);
      } catch (InterruptedException e) {
       e.printStackTrace();
      } 
      
      bll.remove(3);
     }
    
    }
    

    Output
     Added element :1
     Added element :2
     Added element :3
     Added element :4
     Added element :5
     Removed element :3
     Added element :6


    Important points to remember

    1. You should be very careful to make sure that you are releasing after acquire. You can miss it due to programming error or any exception. 
    2. Long held lock can cause starvation 
    3. Method, release doesn't have to be called by the same thread which called acquire. This is an important property that we don't have with normal mutex in Java.
    4. You can increase number of permits at runtime (you should be careful though). This is because number of permits in a semaphore is not fixed, and call to release will always increase the number of permits, even if no corresponding acquire call was made. 

    ---
    keep coding !!!