Showing posts with label solrj. Show all posts
Showing posts with label solrj. Show all posts

Saturday, March 21, 2015

Using CursorMark for deep paging in Solr

Have you ever had to search through 100 million XML blobs to identify which XML blobs had some specific data (or, in my case, was missing some data)? It takes forever, and it's not very fun. I had a situation at work where it looked like I was going to have to do just that. Fortunately, we have the same data (more or less) in Solr cores. I just needed to do a NOT query on a specific field.

The Solr documentation suggests that you use a CursorMark parameter in your query if you need to page through more than a thousand records. 

The following is an example of using CursorMark using SolrJ. The query will return all documents that do not have data for targetField. It then loops through the matches using the cursorMark parameter as a way to tell solr where to retrieve the next set of matches. One of the conditions of using cursorMark to page through results is that you need to sort on a unique field. The schema we are using has a unique field named "uid". 

Something to note is that I had to explicitly set "timeAllowed" to 0 to say that I didn't want any time restrictions on the query.

Yonik Seeley made an interesting point in his Solr 'n Stuff blog about using cursorMark. You can change the returned fields, the facet fields, and number of rows returned, for a specific cursorMark since the cursorMark contains the state of the current search - the state isn't stored server-side. This makes it very easy for a client to allow a user to vary their experience while they page through results. This feels very similar to how you page through results using MySQL and other databases.

import org.apache.solr.client.solrj.SolrQuery;
import org.apache.solr.client.solrj.SolrServer;
import org.apache.solr.client.solrj.SolrServerException;
import org.apache.solr.client.solrj.impl.HttpSolrServer;
import org.apache.solr.client.solrj.response.QueryResponse;
import org.apache.solr.common.SolrDocument;
import org.apache.solr.common.SolrDocumentList;
import org.apache.solr.common.params.CursorMarkParams;
import java.io.*;
import java.util.zip.ZipEntry;
import java.util.zip.ZipOutputStream;

public class FindMissingData {

  private static final int MAX_ALLOWED_ROWS=1000;
  private static final String URL_PREFIX="http://somesolrserver:8983/solr/thecore";
  private static final String ZIP_ENTRY_NAME="TheData.txt";
  private static final String UNIQUE_ID_FIELD="uid";
  private static final String SOLR_QUERY="NOT targetField:[* TO *]";

  public static void main(String[] args) {

    if (args.length != 2) {
      System.out.println("java FindMissingData <rows per batch - max is 1000> <output zip file>");
    }

    int maxRows = Integer.parseInt(args[0]);
    String outputFile = args[1];

    // Delete zip file if it already exists - going to recreate it anyway
    File f = new File(outputFile);
    if(f.exists() && !f.isDirectory()) {
      f.delete();
    }

    FindMissingData.writeIdsForMissingData(outputFile, maxRows);
  }

  public static void writeIdsForMissingData(String outputFile, int maxRowCount) {

    if (maxRowCount > MAX_ALLOWED_ROWS) 
      maxRowCount = MAX_ALLOWED_ROWS;

    FileOutputStream fos = null;
    ZipOutputStream zos = null;

    try {
      fos = new FileOutputStream(outputFile, true);
      zos = new ZipOutputStream(fos);
      zos.setLevel(9);

      queryForMissingData(maxRowCount, zos);
    } catch (Exception e) {
      e.printStackTrace();
    } finally {
      if (zos != null) {
        try {
          zos.flush();
          zos.close();
        } catch (Exception e) {}
      }
      if (fos != null) {
        try {
          fos.flush();
          fos.close();
        } catch (Exception e) {}
      }
    }
  }

  private static void queryForMissingData(int maxRowCount, ZipOutputStream zos) throws IOException {
    ZipEntry zipEntry = new ZipEntry(ZIP_ENTRY_NAME);
    zos.putNextEntry(zipEntry);


    SolrServer server = new HttpSolrServer(URL_PREFIX);

    SolrQuery q = new SolrQuery(SOLR_QUERY);
    q.setFields(UNIQUE_ID_FIELD);
    q.setRows(maxRowCount);
    q.setSort(SolrQuery.SortClause.desc(UNIQUE_ID_FIELD));

    // You can't use "TimeAllowed" with "CursorMark"
    // The documentation says "Values <= 0 mean 
    // no time restriction", so setting to 0.
    q.setTimeAllowed(0);

    String cursorMark = CursorMarkParams.CURSOR_MARK_START;
    boolean done = false;
    QueryResponse rsp = null;
       
    while (!done) {
      q.set(CursorMarkParams.CURSOR_MARK_PARAM, cursorMark);
      try {
        rsp = server.query(q);
      } catch (SolrServerException e) {
        e.printStackTrace();
        return;
      }

      writeOutUniqueIds(rsp, zos);

      String nextCursorMark = rsp.getNextCursorMark();
      if (cursorMark.equals(nextCursorMark)) {
        done = true;
      } else {
        cursorMark = nextCursorMark;
      }
    }
  }

  private static void writeOutUniqueIds(QueryResponse rsp, ZipOutputStream zos) throws IOException {
    SolrDocumentList docs = rsp.getResults();

    for(SolrDocument doc : docs) {
      zos.write(
        String.format("%s%n",
          doc.get("uid").toString()).getBytes()
      );
    }
  }
}

Wednesday, November 20, 2013

Index POJOs into Solr using SolrJ and SolrServer.addBeans()

SolrJ makes indexing and querying data from solr instances very easy.  One way that SolrJ helps makes indexing data easier is that you can create a POJO (Plain Old Java Object) with annotations that will map the fields to your schema.

Note: I used Solr 4.4 (and SolrJ 4.4) when creating this example.


I created a POJO named SampleDoc that maps a number of fields to the example solr schema:

package org.codecognition;

import org.apache.solr.client.solrj.beans.Field;

public class SampleDoc {
    @Field("id")
    private String id;

    @Field("sku")
    private String sku;

    @Field("name")
    private String name;

    @Field("title")
    private String title;

    @Field("price")
    private float price;

    @Field("inStock")
    private boolean inStock;

    public String getId() { return id; }

    public void setId(String id) { this.id = id; }

    public String getSku() { return sku; }

    public void setSku(String sku) { this.sku = sku; }

    public String getName() { return name; }

    public void setName(String name) { this.name = name; }

    public String getTitle() { return title; }

    public void setTitle(String title) { this.title = title; }

    public float getPrice() { return price; }

    public void setPrice(float price) { this.price = price; }

    public boolean isInStock() { return inStock; }

    public void setInStock(boolean inStock) { this.inStock = inStock; }
}


The POJO can then be indexed using the HttpSolrServer.addBeans() method:

package org.codecognition;

import org.apache.solr.client.solrj.SolrServerException;
import org.apache.solr.client.solrj.impl.HttpSolrServer;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;
import java.util.Random;
import java.util.UUID;

public class SolrJWithAddBeansExample {

    public static void main(String[] args) {
        try {
            HttpSolrServer server = new HttpSolrServer("http://localhost:8983/solr");

            server.addBeans(makeTestRecords(100));

            server.commit();
        } catch (SolrServerException | IOException e) {
            e.printStackTrace();
        }
    }

    private static List<SampleDoc> makeTestRecords(int numRecords) {

        List<SampleDoc> sampleDocs = new ArrayList<SampleDoc>(numRecords);
        Random random = new Random();

        for (int i = 0; i < numRecords; i++) {
            SampleDoc sampleDoc = new SampleDoc();

            sampleDoc.setId((UUID.randomUUID().toString()));
            sampleDoc.setInStock(true);
            sampleDoc.setName(String.format("test%s", random.nextInt(10000)));
            sampleDoc.setPrice(random.nextFloat());
            sampleDoc.setSku("somesku");
            sampleDoc.setTitle(String.format("test title %s", random.nextInt(10000)));
            sampleDocs.add(sampleDoc);
        }

        return sampleDocs;
    }
}

You can find the code here.

Friday, November 23, 2012

Solr - Indexing Data Using SolrJ and addBeans

So far it looks like indexing data using SolrJ is considerably slower than indexing data using the update handler and a local CSV file.  It took about 36 to 40 seconds to index 100000 documents using SolrServer.addBeans() compared to about 17 to 18 seconds using the update handler and a local CSV file.

The code using SolrJ, listed below, was running on the same machine as Solr.

public static void IndexBeanValues(List testRecords) 
    throws IOException, SolrServerException {

    HttpSolrServer server = new HttpSolrServer("http://localhost:8983/solr");
    server.addBeans(testRecords);
    server.commit();
}

I tried passing in an instance to SolrServer, but it didn't make any noticeable difference for timing.  It might make more of a difference instantiating a new instance of SolrServer for each batch if the Java code using SolrJ is running on a different machine than the Solr server being targeted.

Refer to this post for a more detailed code example using SolrJ and addBeans.

Solr - Indexing Data Using SolrJ

I think I found one of the slowest ways possible to index data into Solr.  I'm looking into various ways to index data into Solr:
  • indexing text files local to the server that Solr is running on using the update handler
  • indexing data using an app using SolrJ that is running on the same server as Solr
  • indexing data using an app using SolrJ that is on a different machine on the same network that the Solr server is on
I was able to index 100000 items of data into Solr using the update handler to process a CSV file in about 17 to 18 seconds.  Next I tried indexing the same data using SolrJ.  It took about 6 minutes!  I'm sure that the reason it took so long is the way that I wrote the method to index the data.  

The method looks like this:

    public static void IndexValues(TestRecord[] testRecords) 
        throws IOException, SolrServerException {

        HttpSolrServer server = new HttpSolrServer("http://localhost:8983/solr");
        for(int i = 0; i < testRecords.length; ++i) {
            SolrInputDocument doc = new SolrInputDocument();
            doc.addField("id", testRecords[i].getId());
            for (Integer value : testRecords[i].getLookupIds()) {
                doc.addField("lookupids", value);
            }
            server.add(doc);
            if(i%100==0) server.commit();  // periodically flush
        }
        server.commit();

    }

I'll have to try something similar, but using beans.  It seems like it could be a bit faster if I used the addBeans method to add multiple documents at once.

Tuesday, November 20, 2012

Solr - Querying with SolrJ



I added a method, cleverly named FindIds, to my test code that will find the unique IDs by doing a search on lookup IDs that are in a range from 0 to 10000.  The query string looks like this:

"lookupids:[0 TO 10000]"

Even though the query is incredibly simple, I used the Solr admin page to test the query first.  That way I could know what to expect to see from SolrJ.  If I got different results from SolrJ, then I would know I that I would need to investigate to see why there was a difference.

Here is the code used:

public static void main(String[] args)  {
  ArrayList<String> ids = SolrIndexer.FindIds("lookupids:[0 TO 10000]");

  for (String id : ids) {
    System.out.printf("%s%n", id);
  }
}

public static ArrayList<String> FindIds(String searchString) {
  ArrayList<String> ids = new ArrayList<String>();
  int startPosition = 0;

  try {

    SolrServer server = new HttpSolrServer("http://localhost:8983/solr");
    SolrQuery query = new SolrQuery();
    query.setQuery(searchString);
    query.setStart(startPosition);
    query.setRows(20);
    QueryResponse response = server.query(query);
    SolrDocumentList docs = response.getResults();
    
    while (docs.size() > 0) {
      startPosition += 20;

      for(SolrDocument doc : docs) {
        ids.add(doc.get("id").toString());
      }

      query.setStart(startPosition);
      response = server.query(query);
      docs = response.getResults();
    }

  } catch (SolrServerException e) {
    e.printStackTrace();  
  }
  
  return ids;
}

If you don't set the number of rows to be returned using query.setRows(), then the default number of rows to be returned will be used.  The default used in the Solr example config is 10.  

If the query results in nothing being found, then the SolrDocumentList is instantiated with 0 items.  If there is an error with the query string, then an Exception will be thrown.

Something nice to add is a sort field.  ie, query.setSortField("id", SolrQuery.ORDER.asc);  

There are still some areas that I want to explore with SolrJ, so I will be posting more in the near future.

Monday, November 19, 2012

Solr - Indexing data using SolrJ


I'm using Solr at work, so I've been experimenting at home with various ways to index data into Solr.  The latest method I tried using is SolrJ.

The Setup

I have Solr set up on my Windows 7 box - I just downloaded the Solr 4.0 zip from http://www.apache.org/dyn/closer.cgi/lucene/solr/4.0.0.   I'm using IntelliJ Community Edition, so I created a new project and then added references to the necessary jar files by going to the Project Settings | Libraries section.  I clicked the + sign, and picked "Java", and then selected all of the SolrJ related jars.  SolrJ is distributed with Solr, and the related jar files for using SolrJ can be found in %SOLR_HOME%\dist and %SOLR_HOME%\dist\solrj-lib.

The Code

Using SolrJ to index data into Solr is amazingly simple.  The sample code found at solrtutorial.com is almost useable as copy and paste.  The version of Solr that the solrjtutorial site is targeting is for a version older than 4.0.  Everything will work as long as you change the import for CommonsHttpSolrServer to HttpSolrServer.  Of course you will also want to use field names that match your schema, but if you use the example solr instance (and therefore the example schema.xml) as a way to test your code then it will work fine.

First, I updated the %SOLR_HOME%\example\solr\collection1\conf\schema.xml file by adding a multi-valued int parameter called "lookupids".

<field name="lookupids" type="int" indexed="true" stored="true" multiValued="true"/>

I reloaded the core using the Solr admin page (from the Solr admin page click Core Admin, and then click the Reload button.  It should turn green after loading if the schema is valid.) to make sure that I didn't manage to screw up the schema.

Second, I created a new Java project using IntelliJ.  I had the code read a CSV file to populate an array of objects lookup IDs.  I used "lookup" IDs for no particular reason other than I thought it made as much sense as using any other arbitrary property name to search on.

After the code loads an array of Widgets, I had the code call a method called IndexValues.  Here is the mostly copy and paste code from the solrtutorial site:

public static void IndexValues(String solrDocId, List<Widget> widgets) throws 
    IOException, SolrServerException {

    HttpSolrServer server = new HttpSolrServer("http://localhost:8983/solr");
    for(Widget widget : widgets) {
        SolrInputDocument doc = new SolrInputDocument();
        doc.addField("id", solrDocId);
        for (Integer value : widget.getLookUpIds()) {
            doc.addField("lookupids", value);
        }
        server.add(doc);
        if(i%100==0) server.commit();  // periodically flush
    }
    server.commit();

}


It would be a good idea to have the URL and number of items to index between commits configurable, but I just wanted to get data indexed with as little work as possible.  It was very simple thanks to the solrtutorial.com site.

Also, it might be a good idea to make the HttpSolrServer instantiated with a singleton provider class.  The provider class could have helper methods for pinging the Solr instance, and for doing an optimize.  There might be well known patterns to follow, so I would read the wiki and look for existing examples first before creating a SolrJ based utility.

Let me know if you come across any good practices to follow, or pitfalls to avoid, when using SolrJ.

I'm going to try using SolrJ to do searches next.  Let me know if there is anything in particular I should watch out for, or if there is anything that you would like me to write about regarding Solr, SolrJ, etc.