VMWare network and vmxnet3 driver
Phenno
Member Posts: 630
Has anyone experienced issues with NAV on vmware virtual machine using vmxnet3 driver?
I'm currently investigating strange issues with NAS service on NAV2015 instance. I was just watching scheduled job doing delete of invoiced sales order (a lot of them to delete). This has been done out of working hours so both NAV and SQL servers aren't doing much in that period. While looking at processor time I'm seeing that NAV is doing some job for a minute, than it slowes down for several minutes. Than it goes up again, doing some work for a minute, than it slows down again for several minutes.
In period of doing some work it generates lot of network transfers (up to 15Mbps), while when he's doing slowly it generates no more than 355Kbps.
Also, I was checking progress via SQL and can see that in good period it deletes 20 or so sales headers per second, while in slow period it deletes very few.
Since I excluded cpu, memory, disk resource issues on NAV and SQL I'm looking at possible network issues with vmxnet3 network adapter. Once I read artikl mentioning issues with this driver so If anyone had similar experience, please share with us.
Article was here: https://kb.vmware.com/selfservice/microsites/search.do?language=en_US&cmd=displayKC&externalId=2008925
Also, I read similar article mentioning flooding SQL with NAS jobs, here (http://www.dynamics.is/?p=2487), though it applies to NAV2016.
I'm currently investigating strange issues with NAS service on NAV2015 instance. I was just watching scheduled job doing delete of invoiced sales order (a lot of them to delete). This has been done out of working hours so both NAV and SQL servers aren't doing much in that period. While looking at processor time I'm seeing that NAV is doing some job for a minute, than it slowes down for several minutes. Than it goes up again, doing some work for a minute, than it slows down again for several minutes.
In period of doing some work it generates lot of network transfers (up to 15Mbps), while when he's doing slowly it generates no more than 355Kbps.
Also, I was checking progress via SQL and can see that in good period it deletes 20 or so sales headers per second, while in slow period it deletes very few.
Since I excluded cpu, memory, disk resource issues on NAV and SQL I'm looking at possible network issues with vmxnet3 network adapter. Once I read artikl mentioning issues with this driver so If anyone had similar experience, please share with us.
Article was here: https://kb.vmware.com/selfservice/microsites/search.do?language=en_US&cmd=displayKC&externalId=2008925
Also, I read similar article mentioning flooding SQL with NAS jobs, here (http://www.dynamics.is/?p=2487), though it applies to NAV2016.
0
Best Answer
-
Not yet confirmed but it seems that issue was somewhere between these two (guest) machines which were sitting od different vmware hosts. Once they were moved to the same host, everything started to work much faster, hence much less locking issues.
Now it raised my question, how to differentiate async_network_io caused by standard NAV way of receiving data from async_network_io caused by network latency.
Btw. ping from site to site showed 1ms. It looks like SQL packets were investigated differently...5
Answers
-
Have a look at this blog post, Stryk says that it occurs with 2015 & 2016 so it maybe the problem you have...
blog.stryk.info/2016/12/06/navsql-is-completely-stalled-because-of-too-many-cpu-threads/1 -
Yes, I already saw this and simptoms are similar but not the same. I did check workers count agains max workers count and numbers are not close. My max workers count is 544 (6cpu) and working threads are 80, mostly. Also, I do not see complete stall but rather slowdown that cannot be corelated to cpu, memory and/or disks.
Btw. I saw that my instance tops 40 open connections (from NAV to SQL) although there is no such limit set anywhere. Is this fixed limit on NAV instance?0 -
Not yet confirmed but it seems that issue was somewhere between these two (guest) machines which were sitting od different vmware hosts. Once they were moved to the same host, everything started to work much faster, hence much less locking issues.
Now it raised my question, how to differentiate async_network_io caused by standard NAV way of receiving data from async_network_io caused by network latency.
Btw. ping from site to site showed 1ms. It looks like SQL packets were investigated differently...5
Categories
- All Categories
- 75 General
- 75 Announcements
- 66.7K Microsoft Dynamics NAV
- 18.8K NAV Three Tier
- 38.4K NAV/Navision Classic Client
- 3.6K Navision Attain
- 2.4K Navision Financials
- 116 Navision DOS
- 851 Navision e-Commerce
- 1K NAV Tips & Tricks
- 772 NAV Dutch speaking only
- 611 NAV Courses, Exams & Certification
- 2K Microsoft Dynamics-Other
- 1.5K Dynamics AX
- 253 Dynamics CRM
- 103 Dynamics GP
- 6 Dynamics SL
- 1.5K Other
- 991 SQL General
- 383 SQL Performance
- 34 SQL Tips & Tricks
- 28 Design Patterns (General & Best Practices)
- Architectural Patterns
- 9 Design Patterns
- 4 Implementation Patterns
- 53 3rd Party Products, Services & Events
- 1.6K General
- 1K General Chat
- 1.6K Website
- 77 Testing
- 1.2K Download section
- 23 How Tos section
- 249 Feedback
- 12 NAV TechDays 2013 Sessions
- 13 NAV TechDays 2012 Sessions