Showing posts with label c#. Show all posts
Showing posts with label c#. Show all posts

Saturday, October 11, 2014

Testing XSLT changes can be slow if you don't have the right tools

There are probably 50 million tools out there for writing XSLT and viewing the output. I have been in the habit of using a command line app that would take a path to the XSLT file, the source xml file, and a target output file. It worked good enough since I don't usually have to do a lot of XSLT. I just started doing changes on a project that uses XSLT to transform data into XML that can be used to index into a Solr instance, and there are quite a few fields that I need to format in a variety of ways. I was using the simple command line app to view the output, but it was getting a bit tedious.  I was wishing I had something more like IntelliJ - something that would show me the output right away, and it dawned on me that all I needed would be three text boxes and a timer control to get what I wanted.  

Introducing the incredibly named XsltTester. It's pretty basic, and I'm sure that there are existing free tools that do the same thing. However, this took just 30 minutes or so and it was fun to put together.

Here is a link to the project on BitBucket: https://bitbucket.org/lee_wallen/xslttester

Here are a few views of the app:


New icons and buttons.


Example with XML and invalid XSL.

Example with XML and fixed XSL but wrong output.

Example with XML, fixed XSL and fixed output.



Friday, March 22, 2013

Regex in C# to get UTC date...

I recently needed to find a UTC date in a string of text, and I thought it might be handy to pull the various values from the dates that were found by getting the values from the returned Groups for each Match. Groups are identified in your regular expression string by surrounding sections with parentheses. The first Group of every match is the string value that the Match found. Every other group in the Match is identified by the parentheses going from left to right. So, if you have a regular expression that looks like this:
"((\d\d\d\d)-(\d\d)-(\d\d))T((\d\d):(\d\d):(\d\d))Z"
Imagine that someone used the above regular expression on the following string.
"This is a UTC date : 1999-12-31T23:59:59Z. Get ready to party!"
The very first group (after the matched date/time string) will be the entire date, because the first set of parentheses completely wraps the date portion of the string.
"1999-12-31"
The next group would be the year portion of the date, since the next set of parentheses completely wraps the year.
"1999"
That pattern is repeated for the rest of the regular expression string. If no parentheses (groupings) are specified, then there will only be the one group and it will contain the string that the regular expression matched. Here is an example of how to do this in code:
static void Main(string[] args)
{
    string input = "this\tis\ta test 2013-03-21T12:34:56Z\tand\tanother date\t2013-03-21T23:45:01Z";
    string regexString = @"((\d\d\d\d)-(\d\d)-(\d\d))T((\d\d):(\d\d):(\d\d))Z";
    TestRegex(input, regexString);
}

private static void TestRegex(string input, string regexString)
{
    int matchCount = 0;
    foreach (Match match in Regex.Matches(input, regexString))
    {                
        int groupCount = 0;
        foreach (Group group in match.Groups)
        {
            Console.WriteLine("Match {0}, Group {1} : {2}", 
                                matchCount, 
                                groupCount++, 
                                group.Value);    
        }
        matchCount++;
    }
}
Here is the output:
Match 0, Group 0 : 2013-03-21T12:34:56Z
Match 0, Group 1 : 2013-03-21
Match 0, Group 2 : 2013
Match 0, Group 3 : 03
Match 0, Group 4 : 21
Match 0, Group 5 : 12:34:56
Match 0, Group 6 : 12
Match 0, Group 7 : 34
Match 0, Group 8 : 56
Match 1, Group 0 : 2013-03-21T23:45:01Z
Match 1, Group 1 : 2013-03-21
Match 1, Group 2 : 2013
Match 1, Group 3 : 03
Match 1, Group 4 : 21
Match 1, Group 5 : 23:45:01
Match 1, Group 6 : 23
Match 1, Group 7 : 45
Match 1, Group 8 : 01

Wednesday, March 13, 2013

Using Amazon's AWS S3 via the AWS .Net SDK

Amazon's AWS S3 (Simple Storage Service) is incredibly easy to use via the AWS .Net SDK, but depending on your usage of S3 you might have to pay. S3 has a free usage tier option, but the amount of space allowed for use is pretty small by today's standards (5GB). The upside is that even if you end up going outside of the parameters for the free usage tier it is still cheap to use.

Here is some information from Amazon regarding the free usage tier limits for S3:

  • 5 GB of Amazon S3 standard storage, 20,000 Get Requests, and 2,000 Put Requests
  • These free tiers are only available to existing AWS customers who have signed-up for Free Tier after October 20, 2010 and new AWS customers, and are available for 12 months following your AWS sign-up date. When your free usage expires or if your application use exceeds the free usage tiers, you simply pay standard, pay-as-you-go service rates (see each service page for full pricing details). Restrictions apply; see offer terms for more details.


Sign Up To Use AWS

You need to create an account in order to use the Amazon Web Services. Make sure you read the pricing for any service you use so you don't end up with surprise charges. In any case, go to http://aws.amazon.com/ to sign up for an account if you haven't done so already.

Install or Reference the AWS .Net SDK

To start using the AWS .Net SDK to access S3 you will want to either download the SDK from Amazon or use NuGet via Visual Studio. Start Visual Studio (this example is using Visual Studio 2010), and do the following to use NuGet to fetch the AWS SDK:


  • Select the menu item "Tools | Library Package Manager | Manage NuGet Packages For Solution..."
  • Type "AWS" in the "Search Online" search text box
  • Select "AWS SDK for .Net" and click the "Install" button
  • Click "OK" on the "Select Projects" dialog

Create a Project and Use the AWS S3 API

Create a project in Visual Studio, and add the following code:


string key = "theawskeythatyougetwhenyousignuptousetheapis";
string secretKey = "thesecretkeyyougetwhenyousignuptousetheapis";

// create an instance of the S3 TransferUtility using the API key, and the secret key
var tu = new TransferUtility(key, secretKey);

// try listing any buckets you might have
var response = tu.S3Client.ListBuckets();

foreach(var bucket in response.Buckets)
{
   Console.WriteLine("{0} - {1}", bucket.BucketName, bucket.CreationDate);

   // list any objects that might be in the buckets
   var objResponse = tu.S3Client.ListObjects(
      new ListObjectsRequest 
      {
         BucketName = response.Buckets[0].BucketName
      }
   );

   foreach (var s3obj in objResponse.S3Objects)
   {
      Console.WriteLine("\t{0} - {1} - {2} - {3}", s3obj.ETag, s3obj.Key, s3obj.Size, s3obj.StorageClass);
   }
}

// create a new bucket
string bucketName = Guid.NewGuid().ToString();
var bucketResponse = tu.S3Client.PutBucket(new PutBucketRequest
   {
      BucketName = bucketName
   }
);

// add something to the new bucket
tu.S3Client.PutObject(new PutObjectRequest
   {
      BucketName = bucketName,
      AutoCloseStream = true,
      Key = "codecog.png",
      FilePath = "C:\\Temp\\codecog.png"
   }
);

// now list what is in the new bucket (which should only have the one item)
var bucketObjResponse = tu.S3Client.ListObjects(
   new ListObjectsRequest
   {
      BucketName = bucketName
   }
);

foreach (var s3obj in bucketObjResponse.S3Objects)
{
   Console.WriteLine("{0} - {1} - {2} - {3}", s3obj.ETag, s3obj.Key, s3obj.Size, s3obj.StorageClass);
}

Thursday, March 7, 2013

Micro ORM Review - FluentData

Who doesn't love tools that make your life easier? Make a database connection and populate an object in just a few lines of code and one config setting? That's what FluentData can offer. Sign me up! 

Some Features:
  • Supports a wide variety of RDBMs - MS SQL Server, MS SQL Azure, Oracle, MySQL, SQLite, etc.
  • Auto map, or use custom mappers, for your POCOs (or dynamic type).
  • Use SQL, or SQL builders, to insert, update, or delete data.
  • Supports stored procedures.
  • Uses indexed or named parameters.
  • Supports paging.
  • Available as assembly (download the DLL or use NuGET) and as a single source code file.
  • Supports transactions, multiple resultsets, custom return collections, etc.

Pros:
  • Setting up connection strings in a config file, and then passing the key value to a DbContext to establish a connection is such an easy way to do things.  It made it very easy to have generic code point to various databases.  I'm sure that is the intent. Needing to declare a connection object, set the connection string value for the connection object, and then calling the connection object's "Open" method seems undignified. :D It's really not that big of a deal, but I like that it seemed much more straight forward using FluentData.
  • It's very easy to start using FluentData to select, add, update, or delete data from your database.
  • It is easy to use stored procedures.
  • Populating objects from selects, or creating objects and using them to insert new data into your database is almost seamless.
  • Populating more complex objects from selects is fairly easy using custom mapper methods.
  • The exceptions that are thrown by FluentData are actually helpful. The contributors to/creators of FluentData have been very thoughtful in how they return error information.


Cons:
  • I had some slight difficulty setting a parameter for a SQL select when the parameter was used as part of a "like" for a varchar column.  The string value in the SQL looked like this: '@DbName%'.  I worked around the issue by changing the code to use this instead: '@DbName', and then set the value so that it included the %.

I originally thought that I couldn't automap when the resultsets return columns that don't map to properties of the objects (or are missing columns for properties in the target object) without using a custom mapping method. However, there is a way - you can call a method on the DB context to say that automapping failures should be ignored:

Important configurations
  • IgnoreIfAutoMapFails - Calling this prevents automapper from throwing an exception if a column cannot be mapped to a corresponding property due to a name mismatch.
Example Usage:

First, I created a MySQL database to use as a test. I created a database called ormtest, and then created a couple of tables for holding book information:

create table if not exists `authors` (
  `authorid` int not null auto_increment,
  `firstname` varchar(100) not null,
  `middlename` varchar(100),
  `lastname` varchar(100) not null,
  primary key (`authorid` asc)
);

create table if not exists `books` (
 `bookid` int not null auto_increment,
 `title` varchar(200) not null,
 `authorid` int,
 `isbn` varchar(30),
 primary key (`bookid` asc)
);

Next, I created a Visual Studio console app, added an application configuration file, and added a connection string for my database:


  
    
  


Then I created my entity types:

public class Author
{
 public int AuthorID { get; set; }
 public string FirstName { get; set; }
 public string MiddleName { get; set; }
 public string LastName { get; set; }

 public string ToString()
 {
  if (string.IsNullOrEmpty(MiddleName))
  {
   return string.Format("{0} - {1} {2}", AuthorID, FirstName, LastName);
  }
  return string.Format("{0} - {1} {2} {3}", AuthorID, FirstName, MiddleName, LastName);
 }
}
public class Book
{
 public int BookID { get; set; }
 public string Title { get; set; }
 public string ISBN { get; set; }
 public Author Author { get; set; }

 public string ToString()
 {
  if (Author != null)
  {
   return string.Format("{0} - {1} \n\t({2} - {3})", BookID, Title, ISBN, Author.ToString());
  }
  return string.Format("{0} - {1} \n\t({2})", BookID, Title, ISBN);
 }
}
I was then able to populate a list of books by selecting rows from the books table:
public static void PrintBooks()
{
 IDbContext dbcontext = new DbContext().ConnectionStringName("mysql-inventory", new MySqlProvider());
 const string sql = @"select b.bookid, b.title, b.isbn
         from books as b;";
   
 List<Book> books = dbcontext.Sql(sql).QueryMany<Book>();

 Console.WriteLine("Books");
 Console.WriteLine("------------------");
 foreach (Book book in books)
 {
  Console.WriteLine(book.ToString());
 }
}
Unfortunately I wasn't able to select columns from the table that didn't have matching attributes in the entity type. You'll need to create a custom mapping method in order to select extra columns that don't map to any attributes in the entity type. You can also use custom mapping methods to populate entity types that contain attributes of other entity types).

Here is an example:
public static void PrintBooksWithAuthors()
{
 IDbContext dbcontext = new DbContext().ConnectionStringName("mysql-inventory", new MySqlProvider());

 const string sql = @"select b.bookid, b.title, b.isbn, b.authorid, a.firstname, a.middlename, a.lastname, a.authorid 
         from authors as a 
        inner join books as b 
        on b.authorid = a.authorid 
        order by b.title asc, a.lastname asc;";

 var books = new List<Book>();
 dbcontext.Sql(sql).QueryComplexMany<Book>(books, MapComplexBook);

 Console.WriteLine("Books with Authors");
 Console.WriteLine("------------------");
 foreach (Book book in books)
 {
  Console.WriteLine(book.ToString());
 }
}

private static void MapComplexBook(IList<Book> books, IDataReader reader)
{
 var book = new Book
 {
  BookID = reader.GetInt32("BookID"),
  Title = reader.GetString("Title"),
  ISBN = reader.GetString("ISBN"),
  Author = new Author
  {
   AuthorID = reader.GetInt32("AuthorID"),
   FirstName = reader.GetString("FirstName"),
   MiddleName = reader.GetString("MiddleName"),
   LastName = reader.GetString("LastName")
  }
 };
 books.Add(book);
}


And here is an example of an insert, update, and delete:
public static void InsertBook(string title, string ISBN)
{
 IDbContext dbcontext = new DbContext().ConnectionStringName("mysql-inventory", new MySqlProvider());

 Book book = new Book
 {
  Title = title,
  ISBN = ISBN
 };

 book.BookID = dbcontext.Insert("books")
         .Column("Title", book.Title)
         .Column("ISBN", book.ISBN)
         .ExecuteReturnLastId<int>();

 Console.WriteLine("Book ID : {0}", book.BookID);
 
}

public static void UpdateBook(Book book)
{
 IDbContext dbcontext = new DbContext().ConnectionStringName("mysql-inventory", new MySqlProvider());
 book.Title = string.Format("new - {0}", book.Title);

 int rowsAffected = dbcontext.Update("books")
        .Column("Title", book.Title)
        .Where("BookId", book.BookID)
        .Execute();

 Console.WriteLine("{0} rows updated.", rowsAffected);
}

public static void DeleteBook(Book book)
{
 IDbContext dbcontext = new DbContext().ConnectionStringName("mysql-inventory", new MySqlProvider());

 int rowsAffected = dbcontext.Delete("books")
        .Where("BookId", book.BookID)
        .Execute();

 Console.WriteLine("{0} rows deleted.", rowsAffected);
}


Summary:
FluentData has been fairly easy to use and there appears to be a way to accomplish whatever I want to do. If FluentData's documentation had more examples of how to populate entity types (POCOs), then it would have saved me a little bit of time. As it is, the documentation listed multiple ways to accomplish tasks, so it never took long to find a method that would work.

Monday, January 28, 2013

Follow-up for Connection Pooling...

I had mentioned in one of my recent posts that I was told we are not using connection pooling in our C# web service code. I had done some reading, and it seemed pretty clear that connection pooling is enabled by default in the .Net framework, and that you would need to explicitly disable pooling.

I looked through our code and I couldn't find any instance where we set pooling to false for our connection strings. I also wrote some simple code to test if the classes we are using (from Microsoft.Practices.EnterpriseLibrary.Data) might do something odd that would prevent connection pooling, or require us to explicitly state that we want to use connection pooling.

The result of my test code was that the code where I explicitly set "pooling=false" took a longer time to execute than the code without "pooling=false". That shouldn't be surprising, but I was told by multiple people that the code was not using connection pooling. It is understandable why people thought the code was not using pooling, though. For one, there sometimes is a feeling that as code gets old, then the quality of the design degrades. Also, there are lots of coding styles and myths that get passed around as fact. This has been a concrete example for me that everyone should try things out before just accepting something as a fact when it comes to writing software. In addition, it pays to read the documentation!

The test code looks like this:

private const string pooled =
    @"Data Source=someserver;Initial Catalog=dbname;Integrated Security=SSPI;";
private const string notpooled =
    @"Data Source=someserver;Initial Catalog=dbname;Integrated Security=SSPI;Pooling=false;";

public static void Main(string[] args)
{
    int testCount = 10000;
    
    GenericDatabase db = new GenericDatabase(pooled, SqlClientFactory.Instance);
    Console.WriteLine("    pooled : {0}", RunConnectionTest(db, testCount));

    db = new GenericDatabase(notpooled, SqlClientFactory.Instance);
    Console.WriteLine("not pooled : {0}", RunConnectionTest(db, testCount));
}

private static TimeSpan RunConnectionTest(GenericDatabase db, int testCount)
{
    Stopwatch sw = new Stopwatch();
    sw.Start();
    DbConnection conn;
    for (int i = 0; i < testCount; i++)
    {
        conn = db.CreateConnection();
        conn.Open();
        var schema = conn.GetSchema(); // just to make the code do something
        conn.Close();
    }
    sw.Stop();
    return sw.Elapsed;
}


I specifically did not use using statements for the connection. I wanted to mimic what the service code is doing.

The result of the test was that the loop using the connection string that implicitly uses connection pooling finished in just over 1 second. The loop for the connection string that explicitly says to not use connection pooling took about 30 seconds.

I'm pretty happy to know that we don't need to do any updates to our code to make it use connection pooling!

Wednesday, January 23, 2013

Connection pooling and using TCP instead of Named Pipes...

We are experiencing a very odd problem at work. The SQL Servers are needing to be failed over to the backup instances on a regular basis as a work-around for an issue where MS SQL is not appearing to accept connections, or the connections seem to take a while to establish. Microsoft support has been working with my employer since around November, and they seem to not be getting any closer to diagnosing and fixing the problem.

One of the potential symptoms of the issue is that we have intermittent long running web method calls.  The web methods are fairly simple - we make a sproc call, and then populate and return objects with the data.  The SQL DBAs have checked the performance of the sproc calls during the time that the web methods are performing slowly and have assured us that the sproc calls are returning in 200ms or less.

The CPU load on the load balanced web servers seem to be reasonable, so we have been speculating that the slow web method calls are either tied to the mysterious SQL Server connection issue, or due to some unknown issue while trying to connect to the database.

We decided to take a number of steps to try to improve performance.  If the issue is due to the MSC (mysterious SQL connection) issue, then our changes won't make much of a difference.  However, we might find that our code is either exacerbating the MSC issue, or the direct cause of the slow web method calls.  It's doubtful that our code is the sole cause of the MSC issue since multiple non-related SQL servers have experienced the MSC issue.  

Step 1: Add extra performance counters to try to pinpoint where the slow performance issue is occurring

We added performance counters around some of our database connection code, and around the sproc call, to see if there are any performance issues.  The tool we use to view performance counters was missing the new counters we added that would help show how often new connections are being made.  That should hopefully be fixed soon.    

Step 2: Change the connection from using named pipes to TCP

Is it better to use TCP than named pipes?  According to MSDN, the performance of using named pipes and TCP is virtually identical on fast networks.  However, if you are on a slow network, and the SQL Server database is running on a different machine than the application that will be connecting to it, then TCP could give you better performance.  The documentation made it sound as though named pipes might give better performance if the SQL Server instance was on the same machine as the app that is accessing the database.  That is not our configuration, so it appears that there is no benefit to using named pipes.  However, using TCP will be less "talky" (needs less communication back and forth to perform reads) than named pipes, and it can also help streamline communication by using backlog queues.

I hadn't realized that there might be a potential performance benefit of using TCP.  It shows me that I really should read through the documentation for technologies that I use regularly.  It will help me be a lot more thoughtful about how and why I am doing something.  Does using the default protocol (named pipes) work?  Yes. However, we also need to be mindful of the performance our services have when fetching data from our databases.

Step 3: Change the connections to be in a form that will allow connection pooling (if the app isn't already using connection pooling)

I'm pretty new to C# / .Net, so some of what I've been told by my teammates has been taken on faith.  However, I want to learn, so I did some reading.  It appears that the connection pooling should happen regardless of using SqlConnection or DbConnection.  I'll need to look closer at the code, and config files, to find out whether or not our app is disabling pooling, but it might be that we are using pooling, but the following is occurring due to the MSC issue. From an MSDN article on pooling:
When connection pooling is enabled, and if a timeout error or other login error occurs, an exception will be thrown and subsequent connection attempts will fail for the next five seconds, the "blocking period". If the application attempts to connect within the blocking period, the first exception will be thrown again. After the blocking period ends, another connection failure by the application will result in a blocking period that is twice as long as the previous blocking period.  Subsequent failures after a blocking period ends will result in a new blocking periods that is twice as long as the previous blocking period, up to a maximum of five minutes.
So, if the mysterious SQL connection issue is causing a connection to fail, then we should expect exponentially longer waits (up to 5 minutes) for each failure to connect. It would explain why our web method call is taking hardly any time to return usually, and then takes 30+ seconds at other times. 

I just need to prove that we are using connection pooling, or find out why we are not using connection pooling.  I hope that we are using connection pooling, because the connection failure blocking period seems to be a great explanation for what we are seeing. 

Leave a comment, or send an email, if you have any troubleshooting suggestions.

Tuesday, January 22, 2013

QuickCounters goof...

I can't believe how blind I was!  I added a quick counter for a WCF service because I wanted to track how long a certain process takes.  I added code similar to this:


// TODO: LEE DEBUG QuickCounters
var qcRequest = RequestType.Attach(QCCategoryOut, "SomeProcessName", true);
try
{
   SomeCode.Instance.DoStuff(asset);
   qcRequest.SetComplete();
}
catch (Exception ex)
{
   qcRequest.SetAbort();
   throw ex;
}

What happened is that the counter ended up showing some very large value for SomeProcessName's "Request Execution Time (msec)" value.  I couldn't figure out why something that was finishing in what looked like a second was displaying values like 1298793265 for Maximum value.

I spent a good bit of time searching for posts on QuickCounters having invalid values for the execution time, but couldn't find anything.  Finally, I noticed the problem - I forgot the BeginRequest() call!  I don't know if it is a feature of QuickCounters to not throw an exception if you call SetComplete() without first calling BeginRequest, but I suppose it is nice to not have code blowing up in production because you added QuickCounter code incorrectly.

The following code works just fine:


// TODO: LEE DEBUG QuickCounters
var qcRequest = RequestType.Attach(QCCategoryOut, "SomeProcessName", true);
try
{
   qcRequest.BeginRequest();
   SomeCode.Instance.DoStuff(asset);
   qcRequest.SetComplete();
}
catch (Exception ex)
{
   qcRequest.SetAbort();
   throw ex;
}





Tuesday, December 11, 2012

Windows service frustration...

I've started a personal project that uses Solr, since we're using Solr at work and I figured it would be a great way to learn as much as I can about how to setup and use Solr.

My project currently consists of a Windows service that fetches data in one thread, and uses another thread to periodically index the fetched data into Solr.  

This is where the frustration comes in.  I have been debugging the Windows service code as I've been updating the database, or Solr schema, etc.  I'll find something that I want to change in the service code, so I'll uninstall the service, recompile the code, and re-install the service.  Occasionally, when I attempt to uninstall the service, Windows will display a message stating that the service is marked for deletion.  
The specified service has been marked for deletion
I won't be able to install the service until the service is uninstalled, and Windows won't uninstall the service until the machine is rebooted.  I tried the "sc delete <service name>" method for deleting a service, and I get the same message about the service being marked for deletion.  I ended up rebooting, and the service was gone when Windows restarted.  

I did a search for "remove service without rebooting" and I found this blog post.  The post says to uninstall the service with the services window closed.  I'll definitely follow the suggested advice the next time I have this issue.

Has anyone else experienced this problem and found a different way to avoid the problem?

Update: Following the advice from the blog post mentioned above definitely works.  I haven't had an issue as long as I was making sure that the services window was closed when uninstalling.  Perhaps it is coincidence, and I might as well have thrown chicken bones at the computer, but I followed the advice and I haven't seen the problem again.

Update (2013-01-09): I've now seen the "marked for deletion" issue at work while services are being deployed. The services window is most likely not opened on the target machine. My only guess is that there are multiple ways for Windows to decide that the service is in a state of being used, so it decides to mark the service as being marked for deletion to prevent any additional processes from trying to query the service for its state.  For now I will just try to make sure the services window is closed, and suffer through reboots when the service refuses to be uninstalled.

Sunday, November 25, 2012

Solr - Solr .Net Client

I was looking for Solr .Net clients and found SolrNet.  The SolrNet page has links for downloading the binaries, and it also has a link to the git repository.  

I downloaded the SolrNet source code so I could compare the performance of indexing documents using the SolrNet and SolrJ clients, and then built the code.  Next, I created a simple console application and referenced the SolrNet library.  I then created a method that was basically a copy of the Java code I used (while using SolrJ) for reading in a CSV file, and indexing batches of "documents".  The SolrJ version used POJOs (Plain Old Java Objects) with annotations specifying which Solr fields that the properties map to.  The SolrNet version used POCOs (Plain Old CLR Objects) with annotations specificying which Solr fields that the properties map to.

Here is an example of the POCO I used:


public class TestRecord
 {
    [SolrUniqueKey("id")]
    public string ID { get; set; }

    [SolrField("lookupids")]
    public List<int> LookupIDs { get; set; }
}

Here is an example of the code that indexed the values of the test records in batches:

// The solr server was initialized earlier in the code using 
// the following line of code:
// Startup.Init<TestRecord>("http://localhost:8983/solr");

public static void AddValues(List<TestRecord> testRecords)
{
    var solr = ServiceLocator.Current.GetInstance<ISolrOperations<TestRecord>>();
    solr.AddRange(testRecords);
    solr.Commit();
}

It seems to index the data about as fast as the SolrJ code - which isn't terribly surprising.  It appeared that it was slightly slower, but I will need to run multiple tests of varying batch sizes to see how similar or different the results are between SolrNet and SolrJ.

It took ~2.5 minutes to index 100000 documents when using batches of 100 test records, ~50 seconds for batches of 1000 test records, and ~45 seconds for batches of 10000 test records.

SolrNet is very easy to write code for querying against, or indexing into, a Solr index.  I was very pleased!

Friday, November 23, 2012

Reading RSS Feeds with Java and C#


I wanted to read a few RSS feeds using Java or C#.  I started to write my own code for the RSS feed, and quickly realized that using the stubbed code generated by the various XSDs available for RSS is kind of a pain.  I started to rewrite the code to use XML annotations, and that seemed like a bit too much work when it dawned on me that I should have done a search for RSS related code.  That's when I found Rome.

I was able to read an RSS feed with just two lines of code that I copy/pasted from the Rome tutorial page.  Very nice!

Here is a link to the tutorial:
http://wiki.java.net/twiki/bin/view/Javawsxml/Rome05TutorialFeedReader

I was also interested in seeing if there were any libraries for reading RSS feeds using C#.  I found RSS.Net.  http://www.rssdotnet.com/

It was really easy to get started by following the code examples that the author provided.

There is a handy library for you to use if you want to read RSS feeds whether you are using Java or C#.


Tuesday, October 30, 2012

Legacy App Nightmare

I'm working on a project for work that includes an update to a legacy application.  The application was written by contractors about six years ago, and it appears that they weren't given any time or incentive to refactor any of their work.

The legacy app is a Windows form application written using C#.  It probably had a decent design to start - some of the framework of the application makes it easy to add some new functionality to the application.  Other decisions they made are completely perplexing to me, and unfortunately their isn't anyone around to ask questions to of why things were implemented the way they were.

I had to start digging - stepping through the code, and making notes along the way.  One thing that occurred to me is that I never want someone to look at code I've written and feel as frustrated as I've felt when slogging through this particular legacy application code.  I feel like I follow good coding practices, but it is nice to have opportunities to see what things are like when good practices are not followed.

1. Comments are not bad - even in today's "agile" world.  Good coding principles dictate that we should name methods so that we know what the method is doing, but how do we tell people why we implemented the method the way we did it?  We can use comments!  If there is no ambiguity as to why a method is doing what it is doing, then no comment is necessary.  Just remember to consider whether or not the method will make sense to someone that was not involved in the original design when they look at it six years later.

2. Use informative and accurate method names.  This is not a new idea, but it doesn't mean that it is always easy or that it isn't overlooked.  The code I need to modify has a method called ProcessList.  That's great.  If I were to see that in a call stack, then how would I have any idea of what was being done?  I now need to look at the method to see if the contents in the list are being modified, if the contents in the list are keys for some other action that will work on some other bit of data, if the list is being modified, or something else.  Another thing to watch out for when it comes to method names is that your design might shift a bit, and your initial method names might not be quite as informative as they initially were.

3. Use informative variable names.  It is incredibly frustrating to see methods that have variables like this:


List<int> list = new List<int>();

I want to know what the list is being used for.  I don't want to just know that something is a list of ints.  The name might help you recognize a non-fatal logic problem in the code.

Refactoring is important to do on your project prior to releasing to production.  If you wait to refactor as a followup project, then there is a good chance that the time to do the refactoring will never be available.  It makes sense that the time won't become available, because the cost for the refactoring is potentially higher in a followup project/story than it is if you refactor as part of your original project/story.  The reason it might be more expensive is that you could be taking resources away from other projects that could be generating new revenue.  Also, it doesn't take too much time to go by before non-refactored code can become confusing if the why's aren't commented, and the method names are poorly chosen, etc.