In a modern system, random memory access is now the killer. The CPU has cycles to burn. Some of the lessons I've learned on performance recently taught the exact opposite of what was true twenty years ago.
To that end, a modern style of dealing with containers has to be bounded. You need to know where you are, and what your limits are. There's some issue with what operator++() should do when you reach those limits, but the one answer for sure is that it can't just go stepping past and start stomping on whatever's next. To that end, my current libraries have the following concepts for both buffers and containers:
A Range is a beginning and and end.
These never own the storage, they just indicate where it is. This can be used for both the available space to write to, and for the used space containing data within a buffer or container.
A cursor is a range + an iterator.
This is where I've gone a bit past all the other work in C++ containers, but I think this is important. Modern iterators (at least the fun ones), are bidirectional or random access. That means the beginning is as important to keep track of as the end. And copying a cursor should not narrow you to the space you had left, but should allow the copy to head backwards to the front just as easily as progressing to the end.
This also gives us a great data structure for those algorithms like std::rotate() that operate on three iterators.
There have been a lot of people banging around on this for some time. Andrei Alexandrescu wrote a great paper On Iteration that had lots of stuff to say about his implementation of containers and interators for D.
Labels: c++, programming
What I and many others ended up with after diving in there was encumbered lists. This is generally frowned upon by advanced C++ people who point out that there is a standard library with a number of "high quality" containers in them. But we arrived at encumbered lists, because we were making lists of polymorphic objects. The STL containers are all made for uniform types.
To provide a bridge, the boost people threw the PIMPL hammer at it.
As I understand it, the original use of PIMPL was to "hide" full class definition from users of the class (usually at library boundaries), and also to significantly speed up compile times for large systems. Compiling has never been an instantaneous process (sadly). Over my career, I've had to deal with more than one system that took hours to rebuild, so knocking time off is more than an academic curiousity.
Applied to the problem of polymorphic objects in containers, you end up with a PIMPL front-ing object, with a hidden implementation behind an abstract interface. Throw in a smart pointer as well to manage life-cycle and you suddenly have problem of memory coherency, because for your first object allocated you get:
I think there's a better way, but its a shame that no one else has actually built anything that gives you the power of containers, the algorythms, but plays well with polymorphic objects.
Labels: c++, programming, stl
auto ptr= new std::string("Hello World") ;
You have a pretty good idea what you're going to get back from new, and the compiler can figure it out, so here you don't have to say it twice.
But when you go to call some advanced routine in the standard libraries, and you really don't have any idea what you're going to get back but you want to save the value in a structure, then having the examples use auto is really annoying. This page on std::async left me very little idea what to do if I wanted to create a data structure that captured running threads:
auto handle = std::async(std::launch::async,
parallel_sum, mid, end);
int sum = parallel_sum(beg, mid);
return sum + handle.get();
One has to guess at this point that the return values from async are very close to un-nameable, and use templates instead to capture them. This is the very thing that std::function does to wrap up lambdas which are another construct that is practically un-nameable.
Interestingly enough, one place auto saves the day when not even template argument deduction would work is with another C++11 feature called initializer lists:
templated_fn({1, 2, 3}); // FAIL - initializer list has no type, and so T cannot be deduced
auto al = {10, 11, 12}; // special magic for auto
though you can get around the template failure by spelling things out:
templated_fn<std::initializer_list<int>>({1, 2, 3}); // OK
Labels: c++, c++11, programming
He then took the independent C++ library groups to task, not for lack of effort, but for failure of consistency. Even within some of the large libraries (like Boost), pieces don't always well work together.
But unfortunately its worse than that.
Maybe I'm just stuck in my ways, but I worry about a lot of low level details. I worry about memory allocations, I worry about buffer overflow attacks. I try to keep a clear eye on system resource use and even kernel calls. So for instance I don't want to turn std::string loose on a socket and let the inbound communication possibly allocate megabytes of memory. I also want to avoid buffer overflow problems, but with a minimum of overhead, so I almost universally refer to messages as buffers which contain a pointer and a size, and I scan through them using smart pointers which check bounds, but use a minimum of overhead:
buffer_scan pscan( abuffer) ;
while ( pscan ) { putchar( * (pscan ++)) ; }
Other than coming up with a shorter test case for ending, this isn't much different from how you'd scan through data in C. For iterators on lists, I went even closer to historic C syntax:
btl::tlist piter( alist) ;
for ( ; piter ; ++ piter ) { dosomething( * piter ) ; }
But I don't even know what I'm getting myself into when I call into std algorithms. Does std::sort or std::merge allocate and use extra memory? Or are they entirely in place? As I scale up am I going to need O(n), O(log n) or just a few extra bytes? And if these things worry about multi-threading, do I end up acquiring system locks and other things just to get my 20,000 items in order?
And where are basic things like memory mapping a file? Or how do you the equivalent of a printf("%.1f\n", dtmp) using iostreams? The best I could come up with was cout << 0.1 * ( std::floor( 10 * dtmp)) << std::endl. Ugh.
So while I'm happy to read through the documentation and code for Boost, and see if there's anything cool in Poco, I'm more looking for things I can steal, than libraries I can use. And unfortunately that means I'm not helping move the state of C++ libraries forward very fast. But at least for the projects I work on, the interfaces will be solid and consistent.
Labels: c++, c++11, programming
I thought the delegating constructors addition to c++11 was great. Not something I'd need all the time, but it'd probably come in handy once in a while. But as always, there's a hole.
The standard specifies that if a constructor delegates to itself, the program is ill-formed. It also states that in that case no diagnostic is required.So given something like this:
Ref: thenewcpp.wordpress.com
class C
{
public:
C() { }
C(int aval) : C('i') { }
C(char) : C(42) { }
} ;
int main(int N, char ** S)
{
C test(1) ;
return 0 ;
}
gcc 4.8 compiles without warning, and then crashes when you run. Thankfully clang does better:
testerror.cpp:9:13: error: constructor for 'C' creates a delegation cycle [-Wdelegating-ctor-cycles]I think there's a serious chance that gcc could be irrelevant in the coming future.C(char) : C(42) { } ^testerror.cpp:8:3: note: it delegates toC(int aval) : C('i') { } ^testerror.cpp:9:3: note: which delegates toC(char) : C(42) { } ^1 error generated.
Labels: c++, c++11, programming
Unfortunately, C++ constructors challenged the second one:
class SampleBuffer
{
public:
SampleBuffer(int alen) ;
SampleBuffer(char const * astring) ;
protected:
std::unique_ptr m_buffer ;
int m_len ;
} ;
SampleBuffer::SampleBuffer(int alen) : m_len(alen)
{
* m_buffer = new char[alen +1] ;
}
SampleBuffer::SampleBuffer(char const * astring)
{
int tmplen= strlen( astring) ;
* m_buffer = new char[tmplen +1 ];
m_len= tmplen ;
strncpy( * m_buffer, astring, tmplen) ;
}
Either you lived with two copies of the initialization code, or you created a private init()
function which you called from both constructors. Not ideal, but I never worked in a group large
enough that I had to worry about someone trying to call init() other than in the constructor.
Still, it could happen, and that would probably be bad.
In C++11 they added constructor chaining which have shown up in other languages like c sharp and java. So now the constructors can look like this:
SampleBuffer::SampleBuffer(int alen) : m_len(alen)
{
* m_buffer = new char[alen +1] ;
}
SampleBuffer::SampleBuffer(char const * astring) : DRYBuffer(strlen(astring))
{
strncpy( * m_buffer, astring, m_len) ;
}
Already an improvement, especially if you decide to change something like having m_len represent
the size of the buffer including the null terminator (heaven help you tracking down that off by one
error in the original with the two separate code paths).
Obviously this is just an example (only a small step above the other trivial examples out there), but I've done the init() thing before for non-trivial cases, and this will be a handy alternative.
Of course there's the issue of where you can use this, and where you can't. It looks like g++ 4.6 does not support this, but g++ 4.7 on works fine.
Labels: c++, c++11, programming
typedef std::mt19937 rnd_type ;
std::uniform_int_distribution<rnd_type::result_type> urand(0, 400) ;
rnd_type m_rseed ;
m_rseed.seed( aseed) ;
int ival= urand( m_rseed) ;
auto rnd= std::bind(urand, m_rseed) ;
int ival= rnd() ;
std::chrono::microseconds us( 1) ;So what is missing from the pieces of this example? All the chrono stuff seemed to be there, it was the sleep_for call. On my six month old linux box, GCC 4.6 complained that "class std::thread has no member named sleep_for". Cygwin was up to GCC 4.7, and should have supported it, but the package maintainers didn't compile libstdc++ with the option --enable-libstdcxx-time, and so it was conditional'd out, and Visual Studio 2012 was not having any of it either. Strangely enough, a later example with a slightly different variation:
m_thread->sleep_for( urand( m_rseed) * us) ;
std::this_thread::sleep_for( urand( m_rseed) * us) ;worked just fine, but I was still out on two out of three platforms. I switched back to nanosleep() on the linux platforms, but was not finding anything helpful on windows, so I finally threw in a Sleep(0) on windows. That gave me the strangest behavior ever. Unlike the *nix version where I would see a very close count of cycles between two threads which were sleeping random amounts between consuming a global counter; I was getting imbalances on Windows of 40% consistently (like 364 to 136). Turns out that Sleep(0) is nothing at all like Sleep(1) and is its own private hack. So I layered a dithering routine on top of Sleep like this:
int iadd ;and finally got reasonable results from the windows version. With great trepidation, I copied the example for conditional_variable wait_for() straight from here, and except for having to use my own platform specific sleep() routine again, it worked just fine except for one little detail in Visual Studio. The example used a macro for the atomic initializer:
iadd += urand( m_rseed) ;
if ( iadd > 1000 ) { Sleep( 1) ; iadd -= 1000 ; }
else { Sleep( 0) ; }
std::atomic<int> i = ATOMIC_VAR_INIT(0);Visual Studio was having none of it, but changing it to simply i = 0 worked just fine. The release notes for GCC 4.8 say they no longer require platform developers to specify the --enable-libstdcxx-time, so when that comes out in the next few months (?), we'll give it a try again. I have no idea what Visual Studio's plans are. For now you can see my hacked up samples at github, along with other random test code.
Labels: c++, c++11, programming, threads
edata_list::iptr p(mylist) ; while ( p ++ ) { p-> dosomething() ; }
Now they've added regex to C++, and its painful just to look at the examples:
std::tr1::cmatch res;
str = "<h2>Egg prices</h2>";
std::tr1::regex rx("<h(.)>([^<]+)");
std::tr1::regex_search(str.c_str(), res, rx);
std::cout << res[1] << ". " << res[2] << "\n";
Compare to what this would look like in perl:
my $str= "<h2>Egg prices</h2>"; my ($tagstr, $word1, $word2)= $str =~ /<h(.)>([^<]+)/ ; say $word1. ". ". $word2 ;Now, plenty of people have accused perl of being a write only language, but still, someone should have come up with something a little easier to type. Alas, guess we'll just all head over to javascript and get on with implementing solutions.
Labels: c++, c++11, programming
Feb '04
Oops I dropped by satellite.
New Jets create excitement in the air.
The audience is not listening.
Mar '04
Neat chemicals you don't want to mess with.
The Lack of Practise Effect
Apr '04
Scramjets take to the air
Doing dangerous things in the fire.
The Real Way to get a job
May '04
Checking out cool tools (with the kids)
A master geek (Ink Tank flashback)
How to play with your kids