Showing posts with label java. Show all posts
Showing posts with label java. Show all posts

Sunday, August 6, 2017

Create an AWS Lambda using Java...

Here's a quick walk through for creating an AWS lambda using Java. I happen to use IntelliJ with maven, but you can use whatever IDE and package management you prefer to use. You can find a similar walk-through in the online AWS documentation or in the AWS Lambda In Action book.

1. Create an IAM role for the Lambda to use:
  • Click the "Create new role" button.
  • In the "Select role type" section, Click the "Select" button for "AWS Lambda" from the "AWS Service Role" section.
  • Enter the policy name of "AmazonS3FullAccess", click the check box, and click the "Next step" button.
  • Enter a name in the "Role name" text box (for this example, use "hello-lambda-role"), and enter a fitting description in the "Role description" text box. Click the "Create role" button.

2. Create an S3 bucket.

3. Create a Java project for your AWS Lambda code:
  • Using IntelliJ, create a maven project using maven-archetype-quickstart.
  • Add the aws lambda core dependency to the project's pom file:

<dependency>
  <groupId>com.amazonaws</groupId>
  <artifactId>aws-lambda-java-core</artifactId>
  <version>1.1.0</version>
</dependency>
  • Create a class called HelloWorldLambda that implements RequestHandler<String, String>:

public class HelloWorldLambda implements RequestHandler<String, String> {
@Override
public String handleRequest(String input, Context context) {
    String output = "Hello, " + input + "!";
    return output;
}
}
  • Build the project so that the jar is created setting the output jar name to be HelloLambda.jar.


4. Create the lambda in the AWS console:
  • Click on the "Get Started Now" button.
  • Click on the "Blank Function" item.

  • On the "Configure triggers" page, click in the grey dashed square and then select "S3".
    • Select the bucket that you created in step 2.
    • Select the event type "Object Created (All)".
    • Click "Enable trigger".
  • Click the "Next" button. 
  • Enter a name for the lambda like "hello-lambda"
  • Select "Java 8" for the Runtime

  • Click on the "Upload" button and select your HelloLambda.jar.

  • In the "Lambda function handler and role", enter the full package path to your HelloWorldLambda class.
  • Select "Choose an existing role" for the Role section.
  • Select the "hello-lambda-role" that you created in step 1.

  • In the "Tags" section, enter the value "Name" for the key, and "hello-lambda" for the value.

  • In the "Advanced settings", increase the memory to 512 MB. Leave the timeout at 15 seconds.

  • Click the "Create function" button.



5. Test the lambda!

* Go to "Functions" section of the AWS console's Lambda page.
* Select the "hello-lambda" function by clicking the option button.
* Click on the "Actions" drop down, and click on "Test function". The "Input test event" dialg will appear.
* Enter the text "testing", and then click the "Save and test" button.

This will trigger the lambda function, and you'll see the output in the "Execution result" section.


6. Test the lambda with an S3 creation event:

Uploading a text file with a single line of text to your S3 bucket that you created in step 2 will trigger your lambda, and you can see that the lambda is invoked by using the following steps.

  • Go to the AWS Lambda console page, and select the "Functions" section.
  • Click on the "hello-lambda" function. This should take you to the details for your lambda.  
  • Click on the "Monitoring" tab. 


You'll see that you have invocations for both the test run, and the S3 upload. My image shows invocations for multiple file uploads, and multiple tests.


Learn more about AWS Lambdas through AWS Lambda In Action.






Thursday, May 18, 2017

Refactoring Java code using lambdas - Stream.filter

Refactoring Java Code Using Lambdas

I’ve been going through a book, and absolutely loving it - Java 8 in Action: Lambdas, Streams, and functional-style programming.

The thing that I’ve enjoyed the most so far are the notes on when to use certain methods. In the chapter on streams, chapter 3, there is one particular method that really stuck out for me - the filter method. 

The filter method takes a predicate, and returns a stream of all the elements that match the predicate.

The book points out the following:

“Any time you’re looping over some data and checking each element, you might want to think about using the new filter method on Stream.”

Here is an example - the following code will print out even numbers:

List<Integer> numbers = Arrays.asList(1,2,3,4,5,6,7,8);

for (int num : numbers) {
    if (num % 2 == 0) {
       System.out.println(num);
    }
 }

The code above is pretty straight forward, but it could be written like this instead:

List<Integer> numbers = Arrays.asList(1,2,3,4,5,6,7,8);

numbers.stream()
   .filter(n -> n % 2 == 0)
   .forEach(System.out::println);

This might not look like a huge advantage, because there isn’t a lot different between the two. The amount of code is basically the same as well.  However, the benefit of the lambda version is that it can be chained together with other Stream methods.

For example, imagine that you have some data containing user IDs, and some user IDs have a special prefix to indicate a special user type.  You might want to filter out the special users, and then return a list of users without the prefix and with the name in upper case letters.

List<String> userIds = Arrays.asList(“*alice”,”bob”,”*", "charlie","*dana","evelyn","*frank");

return userIds.stream()
   .filter(u -> u.startsWith(“*”) && u.length() > 1)
   .map(u -> u.substring(1).toUpperCase())
   .collect(Collectors.toList());

To do the same thing without using streams would look something like this:

List<String> userIds = Arrays.asList("*alice","bob","", "*", "charlie","*dana","evelyn","*frank");

List<String> specialUsers = new ArrayList<>();

for (String user : userIds) {
   if (user.startsWith("*") && user.length() > 1) {
      specialUsers.add(user.substring(1).toUpperCase());
   }
}

return specialUsers;

You can see that it would require that another variable would have to be declared to hold the special users. Kind of a waste. 


I’ll definitely be keeping my eyes open for code that is iterating over lists and inspecting each element!



Sunday, June 14, 2015

Simple Neo4J example...

I've been using a graph database at work named Titan (https://github.com/thinkaurelius/titan). It's open source, and has some pretty nice features such as the ability to choose between a number of different databases for storage (Cassandra DB, HBase, Berkeley DB, etc), and use Solr or Elastic Search for external indexing of data. However, I wanted to try out Neo4J for home project. Here are a few of the things that I learned.


1. Neo4J is incredibly easy to add to your project using Maven. Just add a dependency like this:

<dependency>
    <groupId>org.neo4j</groupId>
    <artifactId>neo4j</artifactId>
    <version>2.2.2</version>
</dependency>

2. I used Neo4J as an embedded database like this:

GraphDatabaseService graphDb;
graphDb = new GraphDatabaseFactory().newEmbeddedDatabase( DB_PATH );

3. The Neo4J documentation recommended that you use try with resource:

try (Transaction tx = graphDb.beginTx()) {
  populateDb(graphDb, someDataToAdd);
  tx.success();
} catch (Exception e) {
  e.printStackTrace();
}

4. Adding nodes and relationships is very straightforward:

// Add nodes using createNode, then 
// use setProperty for each property
someDataToAdd.forEach(d -> {
  Node node = graphDb.createNode();
  node.setProperty("name", d.getName());
  node.setProperty("email", d.getEmail());
  node.setProperty("id", d.getId());
  node.addLabel(DynamicLabel.label("Person"));
});
// Add relationships by calling createRelationshipTo 
// on the source node.
// Also, search for nodes by calling findNodes
Node friend = graphDb.findNodes(
DynamicLabel.label("Person"), "id", id).next();
node.createRelationshipTo(friend, RelTypes.FRIEND);

5. Get all relationships and nodes by using the GlobalGraphOperations:

GlobalGraphOperations.at(graphDb)
.getAllRelationships().forEach(n -> n.delete());
GlobalGraphOperations.at(graphDb)
.getAllNodes().forEach(n -> n.delete());

It's possible to query the graph using cypher - I just found it to be easier and more intuitive to use the Java API.


Saturday, March 21, 2015

Using CursorMark for deep paging in Solr

Have you ever had to search through 100 million XML blobs to identify which XML blobs had some specific data (or, in my case, was missing some data)? It takes forever, and it's not very fun. I had a situation at work where it looked like I was going to have to do just that. Fortunately, we have the same data (more or less) in Solr cores. I just needed to do a NOT query on a specific field.

The Solr documentation suggests that you use a CursorMark parameter in your query if you need to page through more than a thousand records. 

The following is an example of using CursorMark using SolrJ. The query will return all documents that do not have data for targetField. It then loops through the matches using the cursorMark parameter as a way to tell solr where to retrieve the next set of matches. One of the conditions of using cursorMark to page through results is that you need to sort on a unique field. The schema we are using has a unique field named "uid". 

Something to note is that I had to explicitly set "timeAllowed" to 0 to say that I didn't want any time restrictions on the query.

Yonik Seeley made an interesting point in his Solr 'n Stuff blog about using cursorMark. You can change the returned fields, the facet fields, and number of rows returned, for a specific cursorMark since the cursorMark contains the state of the current search - the state isn't stored server-side. This makes it very easy for a client to allow a user to vary their experience while they page through results. This feels very similar to how you page through results using MySQL and other databases.

import org.apache.solr.client.solrj.SolrQuery;
import org.apache.solr.client.solrj.SolrServer;
import org.apache.solr.client.solrj.SolrServerException;
import org.apache.solr.client.solrj.impl.HttpSolrServer;
import org.apache.solr.client.solrj.response.QueryResponse;
import org.apache.solr.common.SolrDocument;
import org.apache.solr.common.SolrDocumentList;
import org.apache.solr.common.params.CursorMarkParams;
import java.io.*;
import java.util.zip.ZipEntry;
import java.util.zip.ZipOutputStream;

public class FindMissingData {

  private static final int MAX_ALLOWED_ROWS=1000;
  private static final String URL_PREFIX="http://somesolrserver:8983/solr/thecore";
  private static final String ZIP_ENTRY_NAME="TheData.txt";
  private static final String UNIQUE_ID_FIELD="uid";
  private static final String SOLR_QUERY="NOT targetField:[* TO *]";

  public static void main(String[] args) {

    if (args.length != 2) {
      System.out.println("java FindMissingData <rows per batch - max is 1000> <output zip file>");
    }

    int maxRows = Integer.parseInt(args[0]);
    String outputFile = args[1];

    // Delete zip file if it already exists - going to recreate it anyway
    File f = new File(outputFile);
    if(f.exists() && !f.isDirectory()) {
      f.delete();
    }

    FindMissingData.writeIdsForMissingData(outputFile, maxRows);
  }

  public static void writeIdsForMissingData(String outputFile, int maxRowCount) {

    if (maxRowCount > MAX_ALLOWED_ROWS) 
      maxRowCount = MAX_ALLOWED_ROWS;

    FileOutputStream fos = null;
    ZipOutputStream zos = null;

    try {
      fos = new FileOutputStream(outputFile, true);
      zos = new ZipOutputStream(fos);
      zos.setLevel(9);

      queryForMissingData(maxRowCount, zos);
    } catch (Exception e) {
      e.printStackTrace();
    } finally {
      if (zos != null) {
        try {
          zos.flush();
          zos.close();
        } catch (Exception e) {}
      }
      if (fos != null) {
        try {
          fos.flush();
          fos.close();
        } catch (Exception e) {}
      }
    }
  }

  private static void queryForMissingData(int maxRowCount, ZipOutputStream zos) throws IOException {
    ZipEntry zipEntry = new ZipEntry(ZIP_ENTRY_NAME);
    zos.putNextEntry(zipEntry);


    SolrServer server = new HttpSolrServer(URL_PREFIX);

    SolrQuery q = new SolrQuery(SOLR_QUERY);
    q.setFields(UNIQUE_ID_FIELD);
    q.setRows(maxRowCount);
    q.setSort(SolrQuery.SortClause.desc(UNIQUE_ID_FIELD));

    // You can't use "TimeAllowed" with "CursorMark"
    // The documentation says "Values <= 0 mean 
    // no time restriction", so setting to 0.
    q.setTimeAllowed(0);

    String cursorMark = CursorMarkParams.CURSOR_MARK_START;
    boolean done = false;
    QueryResponse rsp = null;
       
    while (!done) {
      q.set(CursorMarkParams.CURSOR_MARK_PARAM, cursorMark);
      try {
        rsp = server.query(q);
      } catch (SolrServerException e) {
        e.printStackTrace();
        return;
      }

      writeOutUniqueIds(rsp, zos);

      String nextCursorMark = rsp.getNextCursorMark();
      if (cursorMark.equals(nextCursorMark)) {
        done = true;
      } else {
        cursorMark = nextCursorMark;
      }
    }
  }

  private static void writeOutUniqueIds(QueryResponse rsp, ZipOutputStream zos) throws IOException {
    SolrDocumentList docs = rsp.getResults();

    for(SolrDocument doc : docs) {
      zos.write(
        String.format("%s%n",
          doc.get("uid").toString()).getBytes()
      );
    }
  }
}

Tuesday, March 3, 2015

Using Mockito to help test retry code...

There are times when you want to make sure that your code can retry an operation that might occaisionally fail due to environment issues.  For example, while using AWS SDK to get an object from S3 you might sometimes get an AmazonClientException - perhaps because of an endpoint being temporarily unreachable. You can simulate this situation using Mockito for the AmazonS3Client.

Mockito provides a fluent style way of how your mock objects respond.  Here is an example:

    AmazonS3Client mockClient = mock(AmazonS3Client.class);
    S3Object mockObject = mock(S3Object.class);

    when(mockClient.getObject(anyString(), anyString()))
        .thenThrow(new AmazonClientException("Something bad happened."))
        .thenReturn(mockObject);

This will cause the code to throw an AmazonClientException on the first call to getObject(), but then return the mocked S3Object on the second call to getObject().

This means that you can have happy path tests where the call eventually succeeds, and a test that will show how your code handles all retries failing, by just setting up your mock with the desired configuration of "thenThrow" and "thenReturn" entries.

Friday, August 23, 2013

Using custom matchers for POJOs with Mockito...

I was trying to do the "right" thing by using TDD while working on a project, and I hit what looks to be a common problem of figuring out how to tell mock objects what to return based on whatever the argument value was that was passed in to the mock's method. It is really easy to specify which values to look for if you are dealing with a simple data type, but it isn't as straight forward if the argument is a complex object. The solution I used is the one I found on StackOverflow.

Here is the code I used:
// member of test class
private MyWorker mockMyWorker = mock(MyWorker.class);
private List<String> listOfOutputValues;
private final String theValueIWant = "some test value";

The class I am testing will use a helper class named MyWorker in this example. The class being tested will call the method MyWorker.listOutputValues(SomeObject) to get a list of output values. The SomeObject has an attribute named someValue. I only want the MyWorker class to return the output listing if the someValue matches a specific value.

I added a "when" statement in the @Before section of the test class. The "when" statement uses the custom matcher to check the value of SomeObject.getSomeValue().

@Before
private void Setup() {
 listOfOutputValues = createTestOutputValues();

 when(mockMyWorker.
                listOutputValues(argThat(hasValidAttributeValue()))).
                   thenReturn(listOfOutputValues);
}

When the mockMyWorker.listOutputValues(SomeObject) method is called in the test code, then it will return the list of Strings if SomeObject.getSomeValue() equals theValueIWant. Here is how the custom matcher is defined in the test class:
// private method in test class
private Matcher<SomeObject> hasValidAttributeValue() {
 return new BaseMatcher<SomeObject>() {
  @Override
  public boolean matches(Object o) {
   return ((SomeObject)o).getSomeValue().equals(theValueIWant);
  }

  @Override
  public void describeTo(Description description) {

  }
 };
}

Monday, May 6, 2013

Things I learned while using AWS SQS...

Updated 03-20-2017

Amazon's Simple Queue Service (SQS) provides an easy to use mechanism for sending and receiving messages between various applications/processes. Here are a few things that I learned while using the AWS Java SDK to use SQS.


SQS is not can be FIFO

It used to be that AWS SQS didn't guarantee FIFO ordering. Now you can create a standard queue or a FIFO queue. However, there are some differences to be aware between standard and FIFO queues that are worth pointing out. The differences can be read about here. Here are some of the key differences:


Standard Queues - available in all regions, nearly unlimited transactions per second, messages will be delivered at least once but might be delivered more than once, messages might be delivered out of order.

FIFO Queues - available in US West (Oregon) and US East (Ohio), 300 transactions per second, messages are delivered exactly once, order of messages is preserved (as the queue type suggests).

SQS Free Usage Tier

The SQS free usage tier is determined by the number of requests you make per month.  You can make up to 1 million requests per month.  The current fee is $.50 per million requests after the first million requests. The cost is pretty low, but it would be easy to start racking up millions of requests. Luckily, there are batch operations that can be done, and each batch operation is considered one request.


Short Polling/Long Polling

You can set a time limit to wait when polling queues for messages. Short polling is when you make a request to receive messages without setting the ReceiveMessageWaitTimeSeconds property for the queue. Setting the ReceiveMessageWaitTimeSeconds property to up to 20 seconds (20 seconds is the maximum wait time) will cause your call to wait up to 20 seconds for a message to appear on the queue before returning.  If there is a message on the queue, then the call will return immediately with the message.  The advantage to using long polling is that you will make less requests without receiving messages. 


One thing to remember is that if you have only one thread being used to poll multiple queues, then you will have unnecessary wait times when only some of the queues have messages waiting.  A solution to that problem is to use one thread for each queue being polled.


Something that seemed a bit contradictory is that queues created through the web console have the ReceiveMessageWaitTimeSeconds set to 0 seconds (meaning it is going to use short polling). However, the FAQ mentions that the AWS SDK uses 20 second wait times by default. I created a queue using the AWS SDK, and the wait time was listed as 0 seconds in the web console. I shouldn't have to specifically set the wait time property to 20 seconds if the default wait time is 20 seconds.  Perhaps the documentation just hasn't been updated yet.


Message Size


The message size can be up to 256 KB in size. If you plan on using SQS as a way to manage a data process flow then you might want to consider how easy it is to reach the 256 KB limit.  Avoid putting data into the queue messages.  Instead, use the messages as notifications for work that needs to be done, and include information that identifies which data is ready to be processed. This is especially important to remember since the messages in the queue can be out of order, and you don't want to count on the data embedded in a message as being the latest version of the data. 


Message TTL On Queues


Messages have a default life span of 4 days on queues, but can be set to be kept for 1 minute to 2 weeks. 


Amazon May Delete Unused Queues


Amazon's FAQ mentions that queues may be deleted if no activity has occurred for 30 days.


JARs Used By AWS Java SDK


There are certain jar files that you will need to reference when using the AWS Java SDK.  They are located in the SDKs "third-party" folder. Here are the jar files I referenced while using the SQS APIs:

  • third-party/commons-logging-1.1.1/commons-logging-1.1.1.jar
  • third-party/httpcomponents-client-4.1.1/httpclient-4.1.1.jar
  • third-party/httpcomponents-client-4.1.1/httpcore-4.1.jar

Friday, February 8, 2013

Code puzzler on DZone and optimizations...

Something that I've really enjoyed about DZone.com is that they have recurring themes for certain posts. One recurring set of posts is the Thursday Code Puzzler. Now and then there will be a really interesting solution. For example, one puzzler was to count the number of 1's that occurred in a set of integers. ie, {1, 2, 11, 14} = 4 since the number 1 occurs 4 times in that set. One of the solutions used map reduce to come up with the solution. I thought that was particularly neat.

I decided to try out the recent code puzzler for finding the largest palindrome in a string. Here was my first method:

public static int getLargestPalindrome(String palindromes) {

        String[] palindromeArray = palindromes.split(" ");
        int largestPalindrome = -1;

        for(String potentialPalindrome : palindromeArray) {
            if (potentialPalindrome.equals(new StringBuffer(potentialPalindrome).reverse().toString()) && potentialPalindrome.length() > largestPalindrome)
                largestPalindrome = potentialPalindrome.length();
        }

        return largestPalindrome;

    }

It works, but it felt like cheating to use the StringBuffer.reverse() method. Here is my second method:
    public static int getLargestPalindrome(String palindromes) {

        int largestPalindrome = -1;

        for(String potentialPalindrome : palindromes.split(" ")) {
            int start = 0;
            int end = potentialPalindrome.length() - 1;
            boolean isPalindrome = true;

            while (start <= end) {
                if (potentialPalindrome.charAt(start++) != potentialPalindrome.charAt(end--)) {
                    isPalindrome = false;
                    break;
                }
            }

            if (isPalindrome && potentialPalindrome.length() > largestPalindrome) {
                largestPalindrome = potentialPalindrome.length();
            }

        }

        return largestPalindrome;

    }

It works faster than the first method. I'm sure the performance improvement is mainly due to the first method's creation of new StringBuffer objects (and also new Strings for the reversed value) for each potential palindrome, but accessing locations in an array for half the length (worst case in 2nd version) is bound to be less work than (worst case in 1st version) comparing every character in a string to another string.

Saturday, February 2, 2013

Thread pool example using Java and ExecutorService...

Using thread pools is something that is very easy to implement using Java's ExecutorService. The Java ExecutorService class allows you to specify the number of asynchronous tasks that you want to process. Here is an example of an ExecutorService class being instantiated where numThreads is an integer specifying the number of threads to create for the thread pool:
ExecutorService executorService = Executors.newFixedThreadPool(numThreads);
You pass a runnable in the ExecutorService's execute method like this:
executorService.execute(someRunnableObject);
I created a sample method that uses an ExecutorService to similute working on text files. It will move files from the source path to an archive path unless there is a "lock" file found. The "lock" file is an empty file that is named identically to one of the files that is being "worked" on. The lock file is used to ensure that the same file isn't attempted to be worked on by multiple threads. I made this sample because I figured this might be a nice way to handle indexing data in csv files to a Solr server. Here is the method that does the work (which would be very poorly named if it weren't sample code):
public static void DoWorkOnFiles(String sourcePath, String archivePath, int numThreads) throws IOException {

    Random random = new Random();
    ExecutorService executorService = Executors.newFixedThreadPool(numThreads);

    File sourceFilePath = new File(sourcePath);
    if (sourceFilePath.exists()) {
        Collection<java.io.File> sourceFiles = FileUtils.listFiles(sourceFilePath, new String[]{"txt"}, false);

        for (File sourceFile : sourceFiles) {
            File lockFile = new File(sourceFile.getPath() + ".lock");
            if (!lockFile.exists()) {
                executorService.execute(new SampleFileWorker(sourceFile.getPath(), archivePath, random.nextInt(10000)));
            }
        }
        // This will make the executor accept no new threads
        // and finish all existing threads in the queue
        try {
            executorService.shutdown();
            executorService.awaitTermination(10000, TimeUnit.MILLISECONDS);
        } catch (InterruptedException ignored) {
        }
        System.out.printf("%nFinished all threads.%n");
    }
    else {
        System.out.printf("%s doesn't exist. No work to do.%n", sourceFilePath);
    }
}
The SampleFileWorker class looks like this:
import org.apache.commons.io.FileUtils;
import java.io.File;
import java.io.IOException;

public class SampleFileWorker implements Runnable {

    private final String sourcePath;
    private final String archivePath;
    private final int testDelay;

    public SampleFileWorker(String sourcePath, String archivePath, int testDelay) {
        this.sourcePath = sourcePath;
        this.archivePath = archivePath;
        this.testDelay = testDelay;
    }

    @Override
    public void run() {

        try {
            File lockFile = new File(sourcePath + ".lock");
            if (!lockFile.exists()) {
                lockFile.createNewFile();
            } else {
                return;
            }

            File sourceFile = new File(sourcePath);
            String archiveFilePath = archivePath.concat(File.separator + sourceFile.getName());
            File archiveFile = new File(archiveFilePath);

            System.out.printf("Simulating work on file %s.%n", sourcePath);
            System.out.printf("Starting: %s%n", sourcePath);
            System.out.printf("Delay:    %s%n", testDelay);

            try {
                Thread.sleep(testDelay);
            } catch (InterruptedException ignored) {
            }

            System.out.printf("Done with: %s%n", sourcePath);
            System.out.printf("Archiving %s to %s.%n", sourceFile, archivePath);

            FileUtils.moveFile(sourceFile, archiveFile);
            sourceFile.delete();
            lockFile.delete();

        } catch (IOException e) {
            e.printStackTrace();
        }
    }
}

Saturday, January 26, 2013

Free networking class via Coursera.org...

I signed up for a class on coursera.org.  Well, that's not true - I signed up for 17 classes on coursera.org.  The classes are all free, and they are taught by professors from top notch universities.  I just started the "Introduction to Computer Networks" class.  It's already been incredibly interesting, and I'm very happy I signed up for it.

The class has already started, but if you don't really care about the grade you get out of the class, and you have an interest in learning how computer networks work, then don't hesitate to sign up. 

The lecture videos have been very straight forward so far, and the video length for each of the lecture segments are in easy to digest (short) segments. This lends itself to being easy to watch during a lunch break, or if you have 30 minutes or less available in the evenings. Most videos have been around 15 minutes or less, but I have found it useful to re-watch segments after realizing that I might have missed something important.

We covered information that explains how latency is computed and it made me realize that, even though it is not difficult to understand latency, there is more to latency than I thought. I have always generalized latency to meaning lag.  ie, playing some online game might be choppy for me due to network latency. That might be true, but now I understand what the latency is.

The way I understood latency explained is that latency is the transmission delay (amount of time it takes to put bits of data onto a "wire") plus the propagation delay (amount of time  for the bits to travel the distance from source host to the target host). 

L = latency, M = message length in bits, R = transmission rate, C = speed of light, and D = length / (2/3) C, then:

Transmission delay : M/R
Propagation delay : R = Length / (2/3) C

Therefore:

Latency : L = M/R + D

When you break it down in that way, then it seems pretty clear as to what might be causing high latency if you check out 1) where the target host is that you are connected to, 2) what your bandwidth is for your network connection, and 3) how much data you are transferring.

Something else we covered was the Nyquist Limit which tells us how much data can be sent given a certain bandwidth.  It was a fairly brief description, and the main thing I got out of it was how to use the formula.  :)  

The formula is :

Rate = 2 * Bandwidth * log(base 2)V bits per second

V is the number of signal levels that are being used.

Here is a sample question from the homework:

TV channels are 6 MHz wide. How much data can be sent per second, if four-level digital signals are used? Assume a noiseless channel, and use the Nyquist limit.

Fun stuff! Apparently the class will eventually include some java programming assignments, so I'm looking forward to those.



Friday, January 18, 2013

Playing with Geocoding APIs...

I decided to mess around with some geocoding APIs.  I didn't really care which API I used, but I wanted it to be free and preferably a REST API.  The REST API "want" is based on the idea that the code for processing the results would be similar for all the REST based geocoding services.

I used both the Yahoo Geocoding API (http://developer.yahoo.com/maps/rest/V1/geocode.html) and the Google Geocoding API (https://developers.google.com/maps/documentation/geocoding/).  

The Yahoo Geocoding API is free and allows up to 10000 requests per day.  The downside is that the API is deprecated, and it looks like they want to use a pay for use service.  In other words, it works right now, but who knows how long?

The Google Geocoding API is free and allows up to 2500 requests per day.  If you exceed the 2500 requests per day, then Google says they may block your usage for a while.  If you continue to exceed the 2500 requests per day, then they will block you indefinitely.  There is a response value for the Google Geocoding API that will let you know you have exceeded your query limit.

I wrote an app that implements both the Yahoo and Google geocoding APIs.  It will either take in a single address/place of interest, or read in addresses from an input file.  The input file is expected to have one address or place of interest per line.

For fun I decided to create a utility class that will take an array of Coordinates (a class that holds the latitude, longitude, and location), and calculate the total distance of each of the Coordinates.  

I found the formula for calculating the distance on stackoverflow.com (providing copy/paste solutions for the masses).  

http://stackoverflow.com/questions/3715521/how-can-i-calculate-the-distance-between-two-gps-points-in-java

I was able to find that the distance between the Space Needle, CenturyLink Field, Big Ben, Taj Mahal, and the Grand Canyon is:

Total distance (km): 27492.11
Total distance (mi): 17082.80

Kind of neat!  

I posted the code on github: https://github.com/leewallen/geocodehelper

If you have any ideas for adding on to the code, then please let me know.  It would be nice to make the app useful to someone.  Also, I wouldn't mind suggestions for any refactoring ideas.

Something that might be a fun Android/Solr project would be to do the following:
  • Create a Solr index for holding information about WIFI access points (this would include the latitude and longitude for the location where the WIFI access point was discovered) 
  • Create an Android app that will log latitude/longitude whenever it finds unsecured (or secured) WIFI access points (ie, Starbucks, libraries, etc.)
  • Index the information into a Solr index
  • Create a new Android app, or add onto the first Android app, to allow you to enter an address/place of interest, and then query Solr to get a list of available WIFI access points within a specified radius
If anyone has any related project ideas, then please share!




Friday, November 23, 2012

Reading RSS Feeds with Java and C#


I wanted to read a few RSS feeds using Java or C#.  I started to write my own code for the RSS feed, and quickly realized that using the stubbed code generated by the various XSDs available for RSS is kind of a pain.  I started to rewrite the code to use XML annotations, and that seemed like a bit too much work when it dawned on me that I should have done a search for RSS related code.  That's when I found Rome.

I was able to read an RSS feed with just two lines of code that I copy/pasted from the Rome tutorial page.  Very nice!

Here is a link to the tutorial:
http://wiki.java.net/twiki/bin/view/Javawsxml/Rome05TutorialFeedReader

I was also interested in seeing if there were any libraries for reading RSS feeds using C#.  I found RSS.Net.  http://www.rssdotnet.com/

It was really easy to get started by following the code examples that the author provided.

There is a handy library for you to use if you want to read RSS feeds whether you are using Java or C#.


Solr - Indexing Data Using SolrJ and addBeans

So far it looks like indexing data using SolrJ is considerably slower than indexing data using the update handler and a local CSV file.  It took about 36 to 40 seconds to index 100000 documents using SolrServer.addBeans() compared to about 17 to 18 seconds using the update handler and a local CSV file.

The code using SolrJ, listed below, was running on the same machine as Solr.

public static void IndexBeanValues(List testRecords) 
    throws IOException, SolrServerException {

    HttpSolrServer server = new HttpSolrServer("http://localhost:8983/solr");
    server.addBeans(testRecords);
    server.commit();
}

I tried passing in an instance to SolrServer, but it didn't make any noticeable difference for timing.  It might make more of a difference instantiating a new instance of SolrServer for each batch if the Java code using SolrJ is running on a different machine than the Solr server being targeted.

Refer to this post for a more detailed code example using SolrJ and addBeans.

Solr - Indexing Data Using SolrJ

I think I found one of the slowest ways possible to index data into Solr.  I'm looking into various ways to index data into Solr:
  • indexing text files local to the server that Solr is running on using the update handler
  • indexing data using an app using SolrJ that is running on the same server as Solr
  • indexing data using an app using SolrJ that is on a different machine on the same network that the Solr server is on
I was able to index 100000 items of data into Solr using the update handler to process a CSV file in about 17 to 18 seconds.  Next I tried indexing the same data using SolrJ.  It took about 6 minutes!  I'm sure that the reason it took so long is the way that I wrote the method to index the data.  

The method looks like this:

    public static void IndexValues(TestRecord[] testRecords) 
        throws IOException, SolrServerException {

        HttpSolrServer server = new HttpSolrServer("http://localhost:8983/solr");
        for(int i = 0; i < testRecords.length; ++i) {
            SolrInputDocument doc = new SolrInputDocument();
            doc.addField("id", testRecords[i].getId());
            for (Integer value : testRecords[i].getLookupIds()) {
                doc.addField("lookupids", value);
            }
            server.add(doc);
            if(i%100==0) server.commit();  // periodically flush
        }
        server.commit();

    }

I'll have to try something similar, but using beans.  It seems like it could be a bit faster if I used the addBeans method to add multiple documents at once.

Monday, November 19, 2012

Solr - Indexing data using SolrJ


I'm using Solr at work, so I've been experimenting at home with various ways to index data into Solr.  The latest method I tried using is SolrJ.

The Setup

I have Solr set up on my Windows 7 box - I just downloaded the Solr 4.0 zip from http://www.apache.org/dyn/closer.cgi/lucene/solr/4.0.0.   I'm using IntelliJ Community Edition, so I created a new project and then added references to the necessary jar files by going to the Project Settings | Libraries section.  I clicked the + sign, and picked "Java", and then selected all of the SolrJ related jars.  SolrJ is distributed with Solr, and the related jar files for using SolrJ can be found in %SOLR_HOME%\dist and %SOLR_HOME%\dist\solrj-lib.

The Code

Using SolrJ to index data into Solr is amazingly simple.  The sample code found at solrtutorial.com is almost useable as copy and paste.  The version of Solr that the solrjtutorial site is targeting is for a version older than 4.0.  Everything will work as long as you change the import for CommonsHttpSolrServer to HttpSolrServer.  Of course you will also want to use field names that match your schema, but if you use the example solr instance (and therefore the example schema.xml) as a way to test your code then it will work fine.

First, I updated the %SOLR_HOME%\example\solr\collection1\conf\schema.xml file by adding a multi-valued int parameter called "lookupids".

<field name="lookupids" type="int" indexed="true" stored="true" multiValued="true"/>

I reloaded the core using the Solr admin page (from the Solr admin page click Core Admin, and then click the Reload button.  It should turn green after loading if the schema is valid.) to make sure that I didn't manage to screw up the schema.

Second, I created a new Java project using IntelliJ.  I had the code read a CSV file to populate an array of objects lookup IDs.  I used "lookup" IDs for no particular reason other than I thought it made as much sense as using any other arbitrary property name to search on.

After the code loads an array of Widgets, I had the code call a method called IndexValues.  Here is the mostly copy and paste code from the solrtutorial site:

public static void IndexValues(String solrDocId, List<Widget> widgets) throws 
    IOException, SolrServerException {

    HttpSolrServer server = new HttpSolrServer("http://localhost:8983/solr");
    for(Widget widget : widgets) {
        SolrInputDocument doc = new SolrInputDocument();
        doc.addField("id", solrDocId);
        for (Integer value : widget.getLookUpIds()) {
            doc.addField("lookupids", value);
        }
        server.add(doc);
        if(i%100==0) server.commit();  // periodically flush
    }
    server.commit();

}


It would be a good idea to have the URL and number of items to index between commits configurable, but I just wanted to get data indexed with as little work as possible.  It was very simple thanks to the solrtutorial.com site.

Also, it might be a good idea to make the HttpSolrServer instantiated with a singleton provider class.  The provider class could have helper methods for pinging the Solr instance, and for doing an optimize.  There might be well known patterns to follow, so I would read the wiki and look for existing examples first before creating a SolrJ based utility.

Let me know if you come across any good practices to follow, or pitfalls to avoid, when using SolrJ.

I'm going to try using SolrJ to do searches next.  Let me know if there is anything in particular I should watch out for, or if there is anything that you would like me to write about regarding Solr, SolrJ, etc.